- The Udacity Inferential Statistics course project is described as:
For the Inferential Statistics final project, you are required to perform a detailed analysis on one or more of the provided data sets. With your chosen data set(s) you will come up with a hypothesis which you wish to test. You will then design an experiment to test this hypothesis and choose an appropriate test. For example, you may use t-tests, ANOVA tests, or any other hypothesis test covered in the course. Remember to check the conditions of any test that you choose to use. Once you’ve run your test you will also need to provide visualizations to support your test. These can be any of the visualizations we learned in this course or in Descriptive Statistics (i.e. histogram, box-whisker plot, scatterplot, etc.).
I have picked the Haberman's Survival Data Set for my project...
This report is a presentation of the results of the exploratory data and statistical analyses I conducted. The working code with the same commentary can be found in the accompanying notebook "Inferential Statistics Project.ipynb". The Python libraries numpy, scipy, pandas, matplotlib, pylab and seaborn were used. Some figures will include information about the underlying library or data types used to create them (e.g. pandas, numpy).
http://archive.ics.uci.edu/ml/datasets/Haberman%27s+Survival
(from http://archive.ics.uci.edu/ml/machine-learning-databases/haberman/haberman.names)
The dataset contains cases from a study that was conducted between 1958 and 1970 at the University of Chicago's Billings Hospital on the survival of patients who had undergone surgery for breast cancer.
-
Number of Instances: 306
-
Number of Attributes: 4 (including the class attribute)
-
Attribute Information:
- Age of patient at time of operation (numerical)
- Patient's year of operation (year - 1900, numerical)
- Number of positive axillary nodes detected (numerical)
- Survival status (class attribute)
- 1 = the patient survived 5 years or longer
- 2 = the patient died within 5 year
-
Missing Attribute Values: None
- Sources: Tjen-Sien Lim (limt@stat.wisc.edu), March 4, 1999
- Past Usage:
- Haberman, S. J. (1976).
Generalized Residuals for Log-Linear Models,
Proceedings of the 9th International Biometrics Conference, Boston, pp. 104-122. - Landwehr, J. M., Pregibon, D., and Shoemaker, A. C. (1984),
Graphical Models for Assessing Logistic Regression Models (with discussion),
Journal of the American Statistical Association 79: 61-83. - Lo, W.-D. (1993).
Logistic Regression Trees, PhD thesis,
Department of Statistics, University of Wisconsin, Madison, WI.
- Haberman, S. J. (1976).
The Haberman's Survival Data has been extensively studied, as it provides an example of medical data that can provide insight into patient survival after a medical intervention. It covers outcomes of patients who have had breast surgery over a 12 year period at the University of Chicago's Billings Hospital, and as such it is an sample of the general population of patients undergoing breast surgery.
It is also a very small dataset (only 306 cases), with only a few features (age, year of operation and number of axillary nodes), which are classified based on the patient's surival or not five years after the surgery. As it contains the actual outcomes it is naturally imbalanced - there are thankfully far more patients who survived, but this imbalance is something that must be accounted for in the data and statistical analysis, and especially beyond that - when predictive models are developed to work with it.
In summary, the dataset is a great example of an important area of research and also provides useful challenges in its analysis, and so continues to be studied. I picked it because I am interested in medical research, it was only as I explored it that I understood its value.
- The research question I am asking is whether there is any significant contributor in the data that can be correlated with the survival outcome - that is, what statistical proof exists for any contributing factors to survival?
- The hypothesis will be formulated after an initial data exploration provides insights into the data
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 306 entries, 0 to 305
Data columns (total 4 columns):
age 306 non-null int64
year 306 non-null int64
nodes 306 non-null int64
survival 306 non-null int64
dtypes: int64(4)
memory usage: 9.7 KB
A few sample rows
| age | year | nodes | survival | |
|---|---|---|---|---|
| 22 | 37 | 60 | 15 | 1 |
| 23 | 37 | 63 | 0 | 1 |
| 24 | 38 | 69 | 21 | 2 |
The survival column was converted into a categorical value, using "Yes" for the original value 1 and "No" for the 2
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 306 entries, 0 to 305
Data columns (total 4 columns):
age 306 non-null int64
year 306 non-null int64
nodes 306 non-null int64
survival 306 non-null category
dtypes: category(1), int64(3)
memory usage: 7.7 KB
The sample rows again
| age | year | nodes | survival | |
|---|---|---|---|---|
| 22 | 37 | 60 | 15 | Yes |
| 23 | 37 | 63 | 0 | Yes |
| 24 | 38 | 69 | 21 | No |
The counts of survival vs non-survival categories
| Survival | Count | Percentage |
|---|---|---|
| Yes | 255 | 73.5 |
| No | 81 | 26.5 |
There are almost 3 times as many subjects surviving their surgery after 5 years compared to those who don't.
- Most years have 2 to 3 times as many survivors to non-survivors
- Except 1965 where both groups have almost the same, fairly high, count
- And 1960 and 1961, where survivors are 6 to 7 times as many
-- We could group by some of these years, but ...
- The intuition is that they relate to clinical practices (e.g. oncology team, surgeon),
- therefore - they will not reflect any general property of people having breast cancer surgery
- and so would be useful for assessment of past hospital practices only...
"During the 1950s and 1960s, the hospital's facilities doubled in size. Adding two cancer research centers, Wyler Children's Hospital, two research laboratories and other leading-edge facilities, the University of Chicago Hospitals doubled from five divisions to 10, increasing the faculty by 100 percent between 1961 and 1971. Between 1963 and 1974, the size of the staff grew again."
Survival by Year (Yes/No)
| Year | Yes | No |
|---|---|---|
| 58 | 24 | 12 |
| 59 | 18 | 9 |
| 60 | 24 | 4 |
| 61 | 23 | 3 |
| 62 | 16 | 7 |
| 63 | 22 | 8 |
| 64 | 23 | 8 |
| 65 | 15 | 13 |
| 66 | 22 | 6 |
| 67 | 21 | 4 |
| 68 | 10 | 3 |
| 69 | 7 | 4 |
Survival grouped by 5 years
| Survival | < 1960 | 1960-65 | > 1965 |
|---|---|---|---|
| Yes | 42 | 123 | 60 |
| No | 21 | 43 | 17 |
Survival divided by the year 1962
| Survival | < 1962 | >= 1962 |
|---|---|---|
| Yes | 89 | 136 |
| No | 28 | 53 |
Survival grouped by those operated on in 1960 and 1961, vs all others
- It seems significant that the relative proportion of survival to non-survival favours survival much more in these two years compared with other years
| Survival | 1960-61 | Other |
|---|---|---|
| Yes | 47 | 178 |
| No | 7 | 74 |
- There are quite a few more survivors compared with non-survivors less than or equal to 40 years old
- Most of the remaining subjects are between 41 and 60
- Look at values in groupings, the 40 and under group may be statistically significant for survival chances
| Age / Survival | 30 | 31 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 | 47 | 48 | 49 | 50 | 51 | 52 | 53 | 54 | 55 | 56 | 57 | 58 | 59 | 60 | 61 | 62 | 63 | 64 | 65 | 66 | 67 | 68 | 69 | 70 | 71 | 72 | 73 | 74 | 75 | 76 | 77 | 78 | 83 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Yes | 3 | 2 | 2 | 5 | 2 | 2 | 6 | 9 | 5 | 3 | 7 | 7 | 7 | 4 | 6 | 3 | 8 | 4 | 8 | 10 | 4 | 10 | 5 | 9 | 8 | 5 | 8 | 7 | 7 | 4 | 6 | 4 | 7 | 5 | 6 | 3 | 4 | 2 | 3 | 5 | 1 | 3 | 2 | 1 | 1 | 1 | 1 | 0 | 0 |
| No | 0 | 0 | 0 | 2 | 0 | 0 | 0 | 1 | 1 | 0 | 3 | 2 | 4 | 3 | 3 | 4 | 3 | 3 | 2 | 2 | 2 | 4 | 6 | 4 | 2 | 2 | 3 | 0 | 1 | 2 | 3 | 3 | 1 | 0 | 4 | 2 | 2 | 0 | 1 | 2 | 0 | 1 | 0 | 1 | 0 | 0 | 0 | 1 | 1 |
Survival by Age group
| Survival | <= 40 | 41-50 | 51-60 | 61-70 | > 70 |
|---|---|---|---|---|---|
| Yes | 39 | 64 | 67 | 45 | 10 |
| No | 4 | 29 | 26 | 18 | 4 |
Survival grouped by Age <= 40 and all others
| Survival | <= 40 | > 40 |
|---|---|---|
| Yes | 39 | 186 |
| No | 4 | 77 |
- Note that for survivors most are concentrated in 4 axillary nodes or less
- Also, both groups have a high number where no nodes exist
| Survival | Yes | No |
|---|---|---|
| Nodes | ||
| 0 | 117 | 19 |
| 1 | 33 | 8 |
| 2 | 15 | 5 |
| 3 | 13 | 7 |
| 4 | 10 | 3 |
| 5 | 2 | 4 |
| 6 | 4 | 3 |
| 7 | 5 | 2 |
| 8 | 5 | 2 |
| 9 | 2 | 4 |
| 10 | 2 | 1 |
| 11 | 1 | 3 |
| 12 | 1 | 1 |
| 13 | 1 | 4 |
| 14 | 3 | 1 |
| 15 | 1 | 2 |
| 16 | 1 | 0 |
| 17 | 0 | 1 |
| 18 | 1 | 0 |
| 19 | 1 | 2 |
| 20 | 1 | 1 |
| 21 | 0 | 1 |
| 22 | 2 | 1 |
| 23 | 0 | 3 |
| 24 | 0 | 1 |
| 25 | 1 | 0 |
| 28 | 1 | 0 |
| 30 | 1 | 0 |
| 35 | 0 | 1 |
| 46 | 1 | 0 |
| 52 | 0 | 1 |
- There are a some with no nodes at all, even here there are fairly large numbers of non-survivors
- Therefore examine in particular the following groups
- Three or less vs all others
- Four or less vs all others
- Between 1 and 3 vs all greater than 3
- Between 1 and 4 vs all greater than 4
- See Appendix entry "Examining Nodes groupings" for tables of other nodes analyses...
Grouped by 3
| Survival | Yes | No |
|---|---|---|
| > 3 | 42 | 47 |
| <= 3 | 39 | 178 |
Grouped by 3 percentages
| Survival | Yes % | No % |
|---|---|---|
| > 3 | 52 | 21 |
| <= 3 | 48 | 79 |
Grouped by 4
| Survival | Yes | No |
|---|---|---|
| > 4 | 39 | 37 |
| <= 4 | 42 | 188 |
Grouped by 4 percentages
| Survival | Yes % | No % |
|---|---|---|
| > 4 | 48 | 16 |
| <= 4 | 52 | 84 |
If the subjects having no nodes at all are included then there are similar numbers of non-survivors in each group.
However, the numbers of survivors are concentrated in those with less than 4 or 3 nodes.
Look at the data with the zero nodes subjects as a separate group, and when excluded...
Grouped by 3 with zero nodes as a separate group
| Survival | Yes | No |
|---|---|---|
| 0 | 117 | 19 |
| 1 to 3 | 61 | 20 |
| > 3 | 47 | 42 |
| Survival | Yes % | No % |
|---|---|---|
| 0 | 52 | 23 |
| 1 to 3 | 27 | 25 |
| > 3 | 21 | 52 |
Grouped by 3 with zero nodes group excluded
| Survival | Yes | No |
|---|---|---|
| 1 to 3 | 61 | 20 |
| > 3 | 47 | 42 |
| Survival | Yes % | No % |
|---|---|---|
| 1 to 3 | 56 | 32 |
| > 3 | 44 | 68 |
Grouped by 4 with zero nodes as a separate group
| Survival | Yes | No |
|---|---|---|
| 0 | 117 | 19 |
| 1 to 4 | 71 | 23 |
| > 4 | 37 | 39 |
| Survival | Yes % | No % |
|---|---|---|
| 0 | 52 | 23 |
| 1 to 4 | 32 | 28 |
| > 4 | 16 | 48 |
Grouped by 4 with zero nodes group excluded
| Survival | Yes | No |
|---|---|---|
| 1 to 4 | 71 | 23 |
| > 4 | 37 | 39 |
| Survival | Yes % | No % |
|---|---|---|
| 1 to 4 | 66 | 37 |
| > 4 | 34 | 63 |
If those having no nodes are excluded, and the remaining are grouped around the threshold of having either 3 or 4 nodes or less, then there is a clear pattern of survival vs non-survival. The pattern is clearest at the 4 node threshold - those having between 1 and 4 nodes vs having greater than 4 nodes - the within-group percentages are almost exactly reversed. The zero-nodes group may be best be treated separately, but we will be not be excluding them from the statistical analysis.
Overall statistical summary
Haberman statistics
age year nodes
count 306.0000 306.0000 306.0000
mean 52.4575 62.8529 4.0261
std 10.8035 3.2494 7.1897
min 30.0000 58.0000 0.0000
25% 44.0000 60.0000 0.0000
50% 52.0000 63.0000 1.0000
75% 60.7500 65.7500 4.0000
max 83.0000 69.0000 52.0000
Statistics with zero nodes removed
Haberman statistics with zero nodes removed
age year nodes
count 170.0000 170.0000 170.0000
mean 51.4588 62.6529 7.2471
std 10.3599 3.2983 8.3551
min 30.0000 58.0000 1.0000
25% 44.0000 60.0000 2.0000
50% 52.0000 62.5000 4.0000
75% 57.7500 65.0000 9.7500
max 83.0000 69.0000 52.0000
Survival group statistics
Survival statistics
age year nodes
count 225.0000 225.0000 225.0000
mean 52.0178 62.8622 2.7911
std 11.0122 3.2229 5.8703
min 30.0000 58.0000 0.0000
25% 43.0000 60.0000 0.0000
50% 52.0000 63.0000 0.0000
75% 60.0000 66.0000 3.0000
max 77.0000 69.0000 46.0000
Survival group statistics with zero nodes excluded
Survival statistics with zero nodes excluded
age year nodes
count 108.0000 108.0000 108.0000
mean 49.7870 62.5000 5.8148
std 10.4139 3.2339 7.3753
min 30.0000 58.0000 1.0000
25% 42.0000 60.0000 1.0000
50% 50.0000 62.0000 3.0000
75% 56.0000 65.0000 7.0000
max 77.0000 69.0000 46.0000
Non-Survival group statistics
Non-Survival statistics
age year nodes
count 81.0000 81.0000 81.0000
mean 53.6790 62.8272 7.4568
std 10.1671 3.3421 9.1857
min 34.0000 58.0000 0.0000
25% 46.0000 59.0000 1.0000
50% 53.0000 63.0000 4.0000
75% 61.0000 65.0000 11.0000
max 83.0000 69.0000 52.0000
Non-Survival group statistics with zero nodes excluded
Non-survival statistics with zero nodes excluded
age year nodes
count 62.0000 62.0000 62.0000
mean 54.3710 62.9194 9.7419
std 9.6721 3.4179 9.3825
min 34.0000 58.0000 1.0000
25% 47.2500 60.0000 3.0000
50% 53.0000 63.0000 7.0000
75% 60.7500 65.0000 13.0000
max 83.0000 69.0000 52.0000
There is very little difference in the distributions of age and year of operation per survival status, but the distribution of the number of auxiliary nodes is quite different, with nodes having a much higher mean and standard deviation in the non-surviving group. This is explored in greater detail in the following charts.
Distributions of values in the Haberman dataset
Distributions of values in the surviving subjects only
Distributions of values in the non-surviving subjects only
Distributions of values per Survival status
The pairplots above illustrate relationships between the Survival outcome and the Age, Year, and Nodes values. The surviving subjects are rendered in blue, the non-surviving subjects are in red.
The distribution curves show no significant difference when plotted over Age, although the range of ages in the not-survived group is slightly smaller due to fewer people under the age of 35.
For the Year there is a dip in the not survived count accompanied by an increase in the survived count around the years 1960 and 1961, and the reverse of this effect around 1967. As this data all comes from the University of Chicago's Billings Hospital its possible that a temporary increase of surviving patients could be due to practises in the hospital for this period - even something like a visiting surgeon or the members of the Oncology team could have contributed to a better outcome.
The Nodes count has a very clear difference in distributions between survived and not survived. The remaining analysis focuses on these differences...
Survival group nodes distribution
count 225.000000
mean 2.791111
std 5.870318
min 0.000000
25% 0.000000
50% 0.000000
75% 3.000000
max 46.000000
Name: nodes, dtype: float64
Non-Survival group nodes distribution
count 81.000000
mean 7.456790
std 9.185654
min 0.000000
25% 1.000000
50% 4.000000
75% 11.000000
max 52.000000
Name: nodes, dtype: float64
The distribution figures above highlight the differences in distribution of between nodes counts in the surviving and non-surviving groups. The mean node count for the surviving group is less than 3, with the 75% percentile of the group having 3 nodes, whereas the not-surviving group have a mean node count of 7.5, and the 75% percentile corresponds to 11 nodes, with 50% corresponding to 4 nodes.
There are quite extreme counts in both groups, but more of them are outliers in the survived group, whereas they mostly fit within the standard distribution of the not-survived group - as can be seen in the further distribution plots below. The plots show that the majority of the surviving group have 3 nodes or less, with those that have more than around 7 being considered as outliers. By contrast, the non-surviving group have a far wider distribution and they don't have any outliers until after around 25 nodes.
Kernel Density Estimates per Survival status
Kernel Density Estimates per Survival status, with outliers removed
- 40 rows removed so we can see core distributon differences closely
The Nodes count has a very clear difference in distributions between survived and not survived. The majority (79%) of the nodes counts in the survived group are less than 4 (84% <= 4), whereas the not-survived group have greater nodes counts in the higher range of the population.
The relationship between nodes counts and survival is explored in terms of node count distributions and kernel density esitmates (KDE) charts above. The KDE estimates the probability density function of a continuous random variable, here it shows that it is roughly twice as probable for non-surviving subjects to have 3 or more nodes compared to the surviving group - that is, from around 3 on the nodes x-axis the line for non-survival has roughly twice the amplitude of the surviving group on the y-axis.
The cumulative distribution function (CDF) is the probability that the variable having a value less than or equal to x.
CDF plots of proportions of counts per features
The CDF plots above are for all subjects, the plot for nodes shows a quite different probability pattern.
CDF plots of Nodes per Survival Status
Focusing again on the relationship between the number of nodes and survival, this CDF has plotted the surviving and non-surviving groups separately.
The top line that approaches the y-axis just above 0.8 illustrates the surviving group - that is, the 84% of of the group that have 4 nodes or less (if you drop a line to the x-axis at that point it intersects at around 3 or 4). The bottom half of the CDF for that group heads down towards zero by the time we reach 14 or 15 on the x axis - meaning that the majority of the data exists with less than around 14 nodes.
The line that approaches the y-axis just under 0.6 illustrates the non-surviving group - that is, the 52% of them having 4 nodes or less, the rest having a higher proportion of nodes compared to the surviving group. The bottom half of this line does not apporach zero until around 32 nodes.
The differences between the distributions are easily discernable, in contrast to the equivalent comparisons for Age and year, produced below...
The CDF plot for Age shows that the Survived group starts before the Non-Surviving group, and that there is a divergence between them around the mid to late 40s, but apart from that they are very similar.
The CDF plot for Year show divergence between 1960 and 1961, and again around 1966.
The following 2 plots illustrate the closer relationships of the Age and Year values per suvival groupings, and that they seem unlikely to contribute significantly to the survival outcome.
The Age box plot illustrates that the first percentile of the surviving group is lower than the non-surviving group, and the third percentile of the non-surviving group is a little higher, but there is little difference between the two.
The violin plot shows more clearly that the non-surviving group is more densly distributed around the age of 50 compared to the surviving group.
Distribution of Age per Survival Status
The Year scatter and violin plots illustrate the decrease in non-survival distributions in the early 1960s, where survival was better for a couple of years. It could be worth exploring why survival improved momentarily in these years.
The decrease in distributions later in the 1960s is more gradual in its effect and could a response to better practises, but the surviving group also decreases, because fewer operations were performed towards the end of the study.
Distribution of Year per Survival Status
- Having axillary nodes is a statistically significant indicator of decreased survival chances
-
It's important to know if any axillary nodes at all are statistically signficant
-
It's useful to understand the difference in possible outcomes around a certain number of nodes
The null hypothesis --> That the number of nodes has no effect on survival outcome.
The alternative hypothesis --> that any number of nodes has a negative effect on survival.We will try and establish if having 4 nodes is a significant number compared to others
- Age grouping by less than or equal to 40 vs. over 40 has a statistically significant effect.
-
If proven it can indicate that being younger than around 40 years of age is an important factor for likely survival
The null hypothesis --> That an age less than or equal to 40 has no effect on survival outcome.
The alternative hypothesis --> that an age less than or equal to 40 has a positive effect on survival.
- Having an operation in the years 1960 and 1961 has a statistically significant effect.
-
This could be of interest if we need to consider changes in the treatments around this potentially significant period
The null hypothesis --> That an operation in the years 1960 or 1961 has no effect on survival outcome.
The alternative hypothesis --> that an operation in the years 1960 or 1961 has a positive effect on survival.
The exploratory data analysis (EDA) shows that the distribution of axillary nodes is quite different between the survived and not suvived groups. Even without a test it's obvious from the differing means and standard deviations that there is an effect. Because of the inequality of the data a Two-Sample t-Test Assuming Unequal Variances (Welch's t-Test) will be conducted. It is expected to show a statistically significant difference and prove the proposed alternative hypothesis that the presence of nodes is a significant indicator of a poor likelihood of survival.
In order to clarify what number of nodes are most indicative of a turning point in survival chances we will conduct a series of Chi-Squared tests on a progressing number of nodes, it is hoped that around 3 or 4 nodes will emerge as the most significant number, since that seems evident from the data analysis.
For the other variables we will conduct standard Two-Sample t-Tests Assuming Equal Variances, as the means and standard deviations between the survived and not survived groups are close for these two groups. It is expected that the test will prove the null hypothesis to be correct.
However, we will also conduct a range of Chi-Squared tests for various groupings of these variables, as the EDA has indicated a group of aged 40 and less have a greater survival chance, as does a group who had their operations in the years 1960 and 1961. It is hoped that these groupings will prove to be statistically significant. The latter group of year of operation, if significant, might point us to review hospital practices during that period.
Hypotheis: Having axillary nodes is a statistically significant indicator of decreased survival chances
- The null hypothesis --> That the number of nodes has no effect on survival outcome.
- The alternative hypothesis --> that any number of nodes has a negative effect on survival
Test 1: Welch's t-Test (Two-Sample t-Test Assuming Unequal Variances)
- Chosen because we are treating the Survived and Non-Survived groups as separate groups, and for node counts their variances differ significantly
Test 2: Chi-Squared tests against the groupings of node counts
- Chosen to test the significance of grouping to specific nodes counts
- It's hoped that it will show a statistical effect most significantly at 4 or 3 nodes
Using alpha level of 0.05 for the tests...
P value and statistical significance: The two-tailed P value is less than 0.0001 By conventional criteria, this difference is considered to be extremely statistically significant.
Confidence interval:
The mean of Survived minus Not Survived equals -4.67
95% confidence interval of this difference: From -6.83 to -2.50
Intermediate values used in calculations:
t = 4.2683
df = 104
standard error of difference = 1.093
Explanation: The unequal variance t test is more useful when you think about it as a way to create a confidence interval. Your prime goal is not to ask whether two populations differ, but to quantify how far apart the two means are. The unequal variance t test reports a confidence interval for the difference between two means that is usable even if the standard deviations differ
Data Summary:
Group Survived Not Survived
Mean 2.79 7.46
SD 5.87 9.19
SEM 0.39 1.02
N 225 81
- Chi-Squared tests were conducted for paired groupings of nodes - that is from less than or equal to a number vs greater than the number
- The ranges tested were from 1 to 7
- The highest Chi-Squared value came with 4 nodes
- We evaluated the Chi-Squared value with Cramer's V test - the effects were medium to small, based on the following definitions
Hypotheis: Age grouping by less than or equal to 40 vs. over 40 has a statistically significant effect.
The null hypothesis --> That an age less than or equal to 40 has no effect on survival outcome.
The alternative hypothesis --> that an age less than or equal to 40 has a positive effect on survival.
Test 1: Two-Sample t-Test Assuming Equal Variances
- Chosen because we are treating the Survived and Non-Survived groups as separate groups, and for Age their variances are similar
- It's expected that the t-Test will prove the null hypothesis when applied to the raw data
Test 2: Chi-Squared test against age groups in 10 year increments
- Chosen to test the significance of a standard grouping
- Like the t-Test, it's not expected to demonstrate statistical significance
Test 3: Chi-Squared test against the 40 and under age group vs those over 40
- Chosen to test the significance of the grouping
- It's hoped that it will show a statistical effect
Using alpha level of 0.05 for the tests...
With a t Statistic of 1.187 we do not exceed either of the one or two-tailed t Critical requirements for significance
- The null hypothesis is proven
- Degrees of freedom = 4
The Chi-squared value does not exceed Chi-critical
- The null hypothesis is proven
- This cannot be considered in isolation from the axillary nodes effect
Having an operation in the years 1960 and 1961 has a statistically significant effect.
- This could be of interest if we need to consider changes in the treatments around this potentially significant period
The null hypothesis --> That an operation in the years 1960 or 1961 has no effect on survival outcome.
The alternative hypothesis --> that an operation in the years 1960 or 1961 has a positive effect on survival.
Test 1: Two-Sample t-Test Assuming Equal Variances
- Chosen because we are treating the Survived and Non-Survived groups as separate groups, and for Year their variances are similar, apart from the years 1960 and 1961, and 1966.
- It's expected that the t-Test will prove the null hypothesis when applied to the raw data
Test 2: Chi-Squared tests against years in various groups
- Chosen to test the significance of a standard grouping
- Like the t-Test, it's not expected to demonstrate statistical significance
Test 3: Chi-Squared test against the years 1960 and 1960 vs all others
- Chosen to test the significance of these two years
- It's hoped that it will show a statistical effect
Using alpha level of 0.05 for the tests...
With a t Statistic of -0.083 we do not exceed either of the one or two-tailed t Critical requirements for significance
- The null hypothesis is proven
- Neither of these show significance
- This could be due to hospital staff and/or practises in those years
-
Conducting an exploratory data analysis provided many insights into the data, which enabled informed hypothesis proposals
-
We assessed the variables of axillary nodes, age, and year of operation into the given categories of Survival or Non-Survival after 5 years
-
The hypotheses for the effect of axillary nodes on survival proposed (1) that having any number of axillary nodes was signficant for the outcome of non-survival, and (2) that 4 nodes was a critical number for deciding on the impact of an increasing number of nodes. Both the t-Test and Chi-Squared tests we conducted supported these alternative hypotheses.
-
The hypotheses for age were that (1) an overall assessment age and (2) for age groups would not be significant, but (3) that a grouping into those 40 years and younger would show a signficant trend for survival in that group. A t-Test supported the first proposal for the null hypothesis, Chi-Squared tests supported both the null hypothesis for insignificance of general age groupings and the alternative hypothesis for the significance of the 40 years and younger grouping.
-
The hypotheses for year were that (1) overall year has no significance, and (2) nor do groupings of year, but that (3) the years 1960 and 1961 as a group compared to all other years would be significant. The t-Test and Chi-Squared tests we conducted confirmed the null hypotheses for the first two proposals and the alternative hypothesis for the third proposal
In conclusion, this analysis showed that even a small sample can hold some very interesting and statistically significant data, and a thorough exploratory data analysis is conducted can assist with understanding what points and groupings in the data might be significant.
-
Title: Haberman's Survival Data
-
Sources: (a) Donor: Tjen-Sien Lim (limt@stat.wisc.edu) (b) Date: March 4, 1999
-
Past Usage:
- Haberman, S. J. (1976). Generalized Residuals for Log-Linear Models, Proceedings of the 9th International Biometrics Conference, Boston, pp. 104-122.
- Landwehr, J. M., Pregibon, D., and Shoemaker, A. C. (1984), Graphical Models for Assessing Logistic Regression Models (with discussion), Journal of the American Statistical Association 79: 61-83.
- Lo, W.-D. (1993). Logistic Regression Trees, PhD thesis, Department of Statistics, University of Wisconsin, Madison, WI.
-
Relevant Information: The dataset contains cases from a study that was conducted between 1958 and 1970 at the University of Chicago's Billings Hospital on the survival of patients who had undergone surgery for breast cancer.
-
Number of Instances: 306
-
Number of Attributes: 4 (including the class attribute)
-
Attribute Information:
- Age of patient at time of operation (numerical)
- Patient's year of operation (year - 1900, numerical)
- Number of positive axillary nodes detected (numerical)
- Survival status (class attribute) 1 = the patient survived 5 years or longer 2 = the patient died within 5 year
-
Missing Attribute Values: None
https://docs.google.com/document/d/1C8l2VxDJEQLJFt97G6hkwbCaR05mW_0BOYz865sILWE/pub
https://docs.google.com/document/d/1JpkpmmZGAhyVZVqgQCEbrznOPXt64mKy8vj3nC_TSFk/pub
https://docs.google.com/document/d/1GyabZyEIFxwt-5_udINmVJG440Cg8Sz0OICf67CUVsc/pub
- Some of the exploratory data analysis concepts were informed by reading through these sites
https://towardsdatascience.com/will-habermans-survival-data-set-make-you-diagnose-cancer-8f40b3449673
https://towardsdatascience.com/exploratory-data-analysis-habermans-cancer-survival-dataset-c511255d62cb
https://www.kaggle.com/kernels/scriptcontent/2743116/download
Distributions of values per Survival status, zero nodes excluded
Distributions of Node counts per Survival status, with zero nodes excluded
Distributions of nodes and Kernel Density Estimates per Survival status, zero nodes excluded
Distribution of nodes and Kernel Density Estimates per Survival Status, zero nodes excluded
Kernel Density Estimates per Survival status, zero nodes excluded
Survival group nodes statistics with zero nodes excluded
count 108.000000
mean 5.814815
std 7.375316
min 1.000000
25% 1.000000
50% 3.000000
75% 7.000000
max 46.000000
Name: nodes, dtype: float64
Non-Survival group nodes statistics with zero nodes excluded
count 62.000000
mean 9.741935
std 9.382466
min 1.000000
25% 3.000000
50% 7.000000
75% 13.000000
max 52.000000
Name: nodes, dtype: float64





















