A Tidy Data Structure and Visualisations for Multiple Variable
Correlations and Other Pairwise Scores
Main:11 Pages
5 Figures
Bibliography:2 Pages
4 Tables
Abstract
We provide a pipeline for calculating, managing and visualising correlations and other pairwise scores for numerical and categorical data. We present a uniform interface for calculating a plethora of pairwise scores and a new tidy data structure for managing the results. We also provide new visualisations which simultaneously show multiple and/or grouped pairwise scores. The visualisations are far richer than a traditional heatmap of correlation scores, as they help identify relationships with categorical variables, numeric variable pairs with non-linear associations or those which exhibit Simpson's paradox. These methods are available in our R package bullseye.
View on arXivComments on this paper
