Research Wiki · Methods Lab

Cluster analysis and segmentation techniques

Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks.

Methods family

Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks.

Use methods only after defining the estimand, data structure, assumptions, validation plan and decision consequence.

Agglomerative clustering

Agglomerative clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Divisive clustering

Divisive clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Single linkage

Single linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Complete linkage

Complete linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Average linkage

Average linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Centroid linkage

Centroid linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Median linkage

Median linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Read the full entry →
Ward’s method

Ward’s method is a hierarchical clustering criterion that merges clusters to minimise the increase in within-cluster variance. Results depend on distance definition, scaling and the cut chosen for the dendrogram.

Read the full entry →
K-means clustering

K-means partitions observations into a chosen number of clusters by minimising within-cluster squared distances to centroids. It is sensitive to scale, initialisation, outliers and the selected number of clusters.

K-medoids

K-medoids is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

PAM clustering

PAM clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

CLARA

CLARA is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

CLARANS

CLARANS is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Fuzzy C-means

Fuzzy C-means is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Possibilistic clustering

Possibilistic clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Gaussian-mixture models

Gaussian-mixture models is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Latent-class analysis

Latent-class analysis estimates unobserved categorical groups from patterns in observed categorical indicators. The classes are probabilistic and require checks for fit, stability, interpretation and external usefulness.

Read the full entry →
Latent-profile analysis

Latent-profile analysis is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Read the full entry →
Finite-mixture modelling

Finite-mixture modelling is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Bayesian mixture models

Bayesian mixture models is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

DBSCAN

DBSCAN forms clusters from dense regions separated by sparser areas and can label isolated observations as noise. It can find irregular shapes but is sensitive to neighbourhood and density parameters.

HDBSCAN

HDBSCAN is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

OPTICS

OPTICS is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Spectral clustering

Spectral clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Affinity propagation

Affinity propagation is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Mean-shift clustering

Mean-shift clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Community detection

Community detection is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Louvain clustering

Louvain clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Leiden clustering

Leiden clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Two-step clustering

Two-step clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Self-organising maps

Self-organising maps is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Neural-network clustering

Neural-network clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Co-clustering

Co-clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Biclustering

Biclustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Consensus clustering

Consensus clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Ensemble clustering

Ensemble clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Constrained clustering

Constrained clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Semi-supervised clustering

Semi-supervised clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Longitudinal clustering

Longitudinal clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Sequence clustering

Sequence clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Trajectory clustering

Trajectory clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Elbow method

Elbow method is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Silhouette coefficient

Silhouette coefficient is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Calinski-Harabasz index

Calinski-Harabasz index is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Davies-Bouldin index

Davies-Bouldin index is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Gap statistic

Gap statistic is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Cluster stability

Cluster stability is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Bootstrap validation

Bootstrap validation is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Holdout validation

Holdout validation is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Segment profiling

Segment profiling is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Segment typing algorithms

Segment typing algorithms is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Browse linked method entries

Deep entry

Agglomerative clustering

Agglomerative clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Divisive clustering

Divisive clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Single linkage

Single linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Complete linkage

Complete linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Average linkage

Average linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Centroid linkage

Centroid linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Median linkage

Median linkage is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Deep entry

Ward’s method

Ward’s method is a hierarchical clustering criterion that merges clusters to minimise the increase in within-cluster variance. Results depend on distance definition, scaling and the cut chosen for the dendrogram.

Open method
Method definition

K-means clustering

K-means partitions observations into a chosen number of clusters by minimising within-cluster squared distances to centroids. It is sensitive to scale, initialisation, outliers and the selected number of clusters.

Open method
Method definition

K-medoids

K-medoids is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

PAM clustering

PAM clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

CLARA

CLARA is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

CLARANS

CLARANS is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Fuzzy C-means

Fuzzy C-means is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Possibilistic clustering

Possibilistic clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Gaussian-mixture models

Gaussian-mixture models is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Deep entry

Latent-class analysis

Latent-class analysis estimates unobserved categorical groups from patterns in observed categorical indicators. The classes are probabilistic and require checks for fit, stability, interpretation and external usefulness.

Open method
Deep entry

Latent-profile analysis

Latent-profile analysis is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Finite-mixture modelling

Finite-mixture modelling is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

Bayesian mixture models

Bayesian mixture models is a method within cluster analysis and segmentation techniques. Algorithms that group observations by similarity, followed by stability, interpretability and actionability checks. Its usefulness depends on data structure, assumptions, validation and whether the output answers the intended decision.

Open method
Method definition

DBSCAN

DBSCAN forms clusters from dense regions separated by sparser areas and can label isolated observations as noise. It can find irregular shapes but is sensitive to neighbourhood and density parameters.

Open method
Method definition

HDBSCAN

HDBSCAN is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

OPTICS

OPTICS is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Spectral clustering

Spectral clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Affinity propagation

Affinity propagation is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Mean-shift clustering

Mean-shift clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Community detection

Community detection is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Louvain clustering

Louvain clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Leiden clustering

Leiden clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Two-step clustering

Two-step clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Self-organising maps

Self-organising maps is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Neural-network clustering

Neural-network clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Co-clustering

Co-clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Biclustering

Biclustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Consensus clustering

Consensus clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Ensemble clustering

Ensemble clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Constrained clustering

Constrained clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Semi-supervised clustering

Semi-supervised clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Longitudinal clustering

Longitudinal clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Sequence clustering

Sequence clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Trajectory clustering

Trajectory clustering is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Elbow method

Elbow method is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Silhouette coefficient

Silhouette coefficient is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Calinski-Harabasz index

Calinski-Harabasz index is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Davies-Bouldin index

Davies-Bouldin index is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Gap statistic

Gap statistic is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Cluster stability

Cluster stability is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Bootstrap validation

Bootstrap validation is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Holdout validation

Holdout validation is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Segment profiling

Segment profiling is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method
Method definition

Segment typing algorithms

Segment typing algorithms is a statistical or analytical concept within cluster analysis and segmentation techniques. It should be selected for the data-generating process and decision question rather than because software makes it available.

Open method

Need the method applied to a live decision?

Share the market, customer, product or investment question that needs a defensible evidence design.

Brief Ninth Atlas