In Python, how can I calculate correlation and statistical significance between two arrays of data?

大兔子大兔子 提交于 2019-12-21 04:12:59

问题


I have sets of data with two equally long arrays of data, or I can make an array of two-item entries, and I would like to calculate the correlation and statistical significance represented by the data (which may be tightly correlated, or may have no statistically significant correlation).

I am programming in Python and have scipy and numpy installed. I looked and found Calculating Pearson correlation and significance in Python, but that seems to want the data to be manipulated so it falls into a specified range.

What is the proper way to, I assume, ask scipy or numpy to give me the correlation and statistical significance of two arrays?


回答1:


If you want to calculate the Pearson Correlation Coefficient, then scipy.stats.pearsonr is the way to go; although, the significance is only meaningful for larger data sets. This function does not require the data to be manipulated to fall into a specified range. The value for the correlation falls in the interval [-1,1], perhaps that was the confusion?

If the significance is not terribly important, you can use numpy.corrcoef().

The Mahalanobis distance does take into account the correlation between two arrays, but it provides a distance measure, not a correlation. (Mathematically, the Mahalanobis distance is not a true distance function; nevertheless, it can be used as such in certain contexts to great advantage.)




回答2:


You can use the Mahalanobis distance between these two arrays, which takes into account the correlation between them.

The function is in the scipy package: scipy.spatial.distance.mahalanobis

There's a nice example here




回答3:


scipy.spatial.distance.euclidean()

This gives euclidean distance between 2 points, 2 np arrays, 2 lists, etc

import scipy.spatial.distance as spsd
spsd.euclidean(nparray1, nparray2)

You can find more info here http://docs.scipy.org/doc/scipy/reference/spatial.distance.html



来源:https://stackoverflow.com/questions/11121762/in-python-how-can-i-calculate-correlation-and-statistical-significance-between

易学教程内所有资源均来自网络或用户发布的内容,如有违反法律规定的内容欢迎反馈
该文章没有解决你所遇到的问题?点击提问,说说你的问题,让更多的人一起探讨吧!