Research Wiki · Methods Lab

Classification and prediction

Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data.

Methods family

Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data.

Use methods only after defining the estimand, data structure, assumptions, validation plan and decision consequence.

Linear discriminant analysis

Linear discriminant analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Read the full entry →
Quadratic discriminant analysis

Quadratic discriminant analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Read the full entry →
Logistic classification

Logistic classification is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Naïve Bayes

Naïve Bayes is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
K-nearest neighbours

K-nearest neighbours is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Decision trees

Decision trees is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Classification and regression trees

Classification and regression trees is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Read the full entry →
Random forests

Random forests combine many decorrelated decision trees fitted to resampled data and random predictor subsets. They often predict well, but importance measures and explanations require care.

Read the full entry →
Extra trees

Extra trees is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Gradient boosting

Gradient boosting is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

XGBoost

XGBoost is a regularised gradient-boosted tree algorithm that sequentially improves an ensemble by fitting residual errors. Performance depends on tuning, validation and control of leakage and overfitting.

LightGBM

LightGBM is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

CatBoost

CatBoost is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Support-vector machines

Support-vector machines is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Neural networks

Neural networks is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Deep learning

Deep learning is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Ensemble learning

Ensemble learning is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Stacking

Stacking is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Bagging

Bagging is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Boosting

Boosting is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Calibration models

Calibration models is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Probability calibration

Probability calibration is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

ROC curves

ROC curves is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

AUC

AUC is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Precision-recall analysis

Precision-recall analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Confusion matrices

Confusion matrices is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Sensitivity and specificity

Sensitivity and specificity is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Lift charts

Lift charts is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Gains charts

Gains charts is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Decile analysis

Decile analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Cross-validation

Cross-validation is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Train-test splitting

Train-test splitting is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Overfitting and underfitting

Overfitting and underfitting is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Hyperparameter tuning

Hyperparameter tuning is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Feature selection

Feature selection is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Feature importance

Feature importance is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

SHAP values

SHAP values is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

LIME explanations

LIME explanations is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Browse linked method entries

Deep entry

Linear discriminant analysis

Linear discriminant analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Deep entry

Quadratic discriminant analysis

Quadratic discriminant analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Deep entry

Logistic classification

Logistic classification is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Naïve Bayes

Naïve Bayes is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

K-nearest neighbours

K-nearest neighbours is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Decision trees

Decision trees is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Classification and regression trees

Classification and regression trees is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Deep entry

Random forests

Random forests combine many decorrelated decision trees fitted to resampled data and random predictor subsets. They often predict well, but importance measures and explanations require care.

Open method
Method definition

Extra trees

Extra trees is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Gradient boosting

Gradient boosting is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

XGBoost

XGBoost is a regularised gradient-boosted tree algorithm that sequentially improves an ensemble by fitting residual errors. Performance depends on tuning, validation and control of leakage and overfitting.

Open method
Method definition

LightGBM

LightGBM is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

CatBoost

CatBoost is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Support-vector machines

Support-vector machines is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Neural networks

Neural networks is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Deep learning

Deep learning is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Ensemble learning

Ensemble learning is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Stacking

Stacking is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Bagging

Bagging is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Boosting

Boosting is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Calibration models

Calibration models is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Probability calibration

Probability calibration is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

ROC curves

ROC curves is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

AUC

AUC is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Precision-recall analysis

Precision-recall analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Confusion matrices

Confusion matrices is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Sensitivity and specificity

Sensitivity and specificity is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Lift charts

Lift charts is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Gains charts

Gains charts is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Decile analysis

Decile analysis is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Cross-validation

Cross-validation is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Train-test splitting

Train-test splitting is a method within classification and prediction. Statistical and machine-learning methods used to assign cases to classes or estimate future outcomes, with validation against unseen data. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Overfitting and underfitting

Overfitting and underfitting is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Hyperparameter tuning

Hyperparameter tuning is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Feature selection

Feature selection is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Feature importance

Feature importance is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

SHAP values

SHAP values is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

LIME explanations

LIME explanations is a statistical or analytical concept within classification and prediction. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method

Need the method applied to a live decision?

Share the market, customer, product or investment question that needs a defensible evidence design.

Brief Ninth Atlas