Tuesday, March 17, 2015

Significance and Chi-Squared Testing

Part 1

*note: for the ‘z/t value’ section, the +/- indicator implies that the critical value is both above and below the mean (o) point.  The values that do not have +/- in front of it will only have one critical value, and could be either in the + range or – range, but not both.

2. The data presented to us in this problem is trying to test if there is a difference between the populations of three Invasive insects in Buck County in comparison to estimated population numbers per field in the entire county.  The null hypotheses is that there is no difference between the invasive species population in the 50 fields sampled from Buck county, and estimated values.  The alternative hypotheses is that there is a difference in these populations.     After calculating the Z scores of the sampled bug populations from the 50 selected fields and comparing their placement in comparison to the 1.96 or -1.96 critical value which was derived from using a two tailed test with 95% confidence, it was concluded upon that null hypotheses for each set of insect data should be rejected.  In conclusion, the insect population for the sampled 50 fields in Buck County for some reason or another have elements that cause there to be more Asian-Long Horned Beatles               (z= 2.47)and Emerald Ash Borer Beetles (z= 7.08) , and less Golden Nematodes ( z = -7.76) , than there are in the predictive model for the county.  

3. Comparing the size of all parties that attended a park in the year 1960, and sample group of 25 parties in 1985, we are trying to see if there is no difference in overall group sizes between the time periods (null hypotheses) or to see if there is a difference (alternative hypotheses).  To test the null hypotheses, we compared the t scores of the sample data to the critical values associated with a one tailed test with 95% confidence level.  T scores were used for this data set because the number of observations is below 30.  The t-score of the sample data was 4.92 and the critical value derived from the 95% confidence level was 1.711.  as a result of the t score being higher then the critical value, we reject the null hypotheses.  In conclusion, these results indicate  that if you were to randomly sample an observation from both time periods, the party from 1985 would have a higher chance of having more group members.  

Part 2: Introduction

                For part two of this assignment, students are too chose three variables and compare the prevalence of said variables between southern counties and northern counties within the state of Wisconsin.  The three variables I chose to investigate (all per county) were ATV trail mileage, number of non-residential gun-deer permits, and number of non-residential 15 day fish license.  I chose these variables because Northern Wisconsin often is associated with rustic wildlife and outdoor fun, and we want to see if the patterns in human behavior and attributes of the land coincide with this idea.   The null hypotheses for this situation would be that there is no real difference in these variables between the counties in the north of the state and the counties in the south of the state.  Conversely, the alternative hypotheses is that there is a difference across the geographic space of north and south for the variables selected.  To test the null hypotheses, the data provided will be mapped to show the spatial distribution of the measured variables.  Subsequently, Chi squared tests will be used to either reject or fail reject the null hypotheses for each variable. 
Figure 2: A visual reference for how the northern counties and southern counties
relate to Highway 29


Methods

                The first step in preparing the data for further analysis is to create all the necessary layers in Arc-map.  To begin, we must join the data provided in SCORPARCGIS table provided in an excel spreadsheet, to a shape file of Wisconsin counties.  The join conducted was based off of the county field from both the counties shape file and the SCORP table,  the cardinality of this join was one-to-one and matched for all 72 counties.  The next step was to add 4 fields to the joined tabled; the first added field will delineate weather a county is north of Highway 29 (1 value) or south of Highway 29 (2 value).  The result was an even split of 36 counties in both the north and south portions. The other three fields added were used to classify my selected variables on a scale of 1,2,3,4. The higher the ranking the more of that variable is in that county.   The 1,2,3,4 ranking system was based on what category a county fell into when symbolized into a cloropleth map that was based on a natural breaks, four class classification.  Once all the values were added for all the new fields for all the counties, the next step is to create the maps and cross tab reports necessary to make conclusions about the null hypotheses.
Results
The Three maps created respectively represent the distribution of ATV trails (miles), non-resident gun-deer licenses, and non-resident 15 day fishing licenses  through the counties of the entire state of Wisconsin.  The colored categories  that you can see in in the legend, and the numbers that they are associated with are  the bases for the 4 categories that were discussed in the methods section of this report.  One of the more important thing that these maps convey to us is how for each map, the counties that represented the highest categorical value (dark green - category 4) are almost entirely in the Northern Counties for all variables.  however, other than that observation, these maps do not indicate fully if there truly is a remarkable difference between the two parts of the state.   We must must do further analyses with the Chi-Squared tests to see if their is a remarkable difference in the spatial occurrence of these variables across the north and south divide. 



After conducting the Chi-Square operation on three selected elements, the following charts were produced.  To reject the null hypotheses the Pearson Chi-Square value had to fall outside of the 9.49 critical value, which is associated with a 95% confidence level corresponding with 3 degrees of freedom.

Non- resident 15 day fish license Score =  12.0
ATV trail mileage score = 15.0
Non-resident gun-deer license score = 6.6

In regards to these results the null hypotheses is rejected for both ATV trail mileage and non-resident 15 day fish licenses.  conversely, we fail to reject the Null hypotheses for non-resident gun-deer licenses.  in the tables below, the values to pay attention too in the first box in each section are the ones in the first row. The Chi-Squared Value for each factor is already listed above, the third value when subtracted from 1 gives you the relative percentage which indicates your confidence in that there is a difference between the northern and southern counties.  The second box in each section displays the difference between the expected values, based off of random selection, and the observed values for each ranked category in the north and the south.


1.  Non-resident 15 day fish license

Score = .007



NOTE: 1 = NORTHERN COUNTIES  
              2= SOUTHERN COUNTIES 






 2.  ATV trail miles



















3.  Non-resident gun- deer licenses

















Conclusion

If looking for a sense of  "Up-North" is the goal, we may have found it.  I believe this based off of my results because two out of my three factors ended up showing a much stronger prevalence in the north than in the south.  Suspicions were rising when i created the map that there was a spatial difference, and running the Chi-Square test confirmed my findings that there is a difference between what is observed and what is expected.  In other words, there was something in the north that was creating the conditions for a higher occurrence of 15 day fishing licenses and more ATV trails. when comparing the maps, the fact that  there is a higher prevalence of non residential fishing licenses implies that this is where the most attractive fishing is in the state.  Attractive in the sense of the number of lakes, variety and availability of fish, and even overall surrounding.  People from out of are traveling all the way up to the Northern part of the state would imply that there is a more rustic and natural vibe 'Up-North'.  Not as convincing or telling an argument, but the higher amount of ATV trail miles indicates even more so that the Northern Counties in Wisconsin is the best destination for outdoor and recreation in the state.  Finding out that the number of non-residential deer licenses was more evenly spread out than the other factors was not to my surprise.  Unlike lakes and trails, deer and deer populations have a much higher degree of mobility, and thus are more randomly spread out. eliminating the out-liers of the northwestern counties that border Minnesota, you might even see more hunters in the south than in the north.  The chi squared values were very helpful in confirming that suspicion that was visual provided by the map, and cemented my conclusion that the physical landscape conditions are different in the northern part of the state, leading to the rejection of the null hypotheses.

Thursday, February 26, 2015

Spatial Statistics: Weighted Mean Centeres and Z Scores.

Introduction

The problem presented in this project pertained to trying to decipher if their has been any geographic shift in where tornado's occur in in the states of Oklahoma and small part of Kansas.  The claim that some citizens of these states make, is that the pattern of tornado events hasn't changed geographically , therefore they shouldn't be forced to build a protective shelter if they live an area with few tornadoes.  The state governments believe that it is in the best interest that everyone should have them, simply to be safe.  To seek sense out of this situation and make a more scientific assessment, their should be a statistical analyses of the areas tornadoes by taking into consideration both the locations and magnitude across two separate but aligned time periods, over the same span of space covering parts of Kansas and Oklahoma.

Methods

The data provided incorporated an almost complete set of information that would be necessary to compare tornado patterns of location and magnitude across the state of Oklahoma and a portion of Kansas. There were three files used to run the analyses on the tornado data: two shape files containing the location and width of the tornado, one set from 1995-2006 and the other from 2007-2012. The third file was shape file of all the counties in the area of interest and also contained the count of tornadoes per county, but only for the 2007-2012 period (hence why I said 'an almost complete set of information').

The first set of analyses that was conducted was to find the weighted mean centers for each time periods.  a weighted mean centers averages out the totality of all the x and y coordinates, and divides both by the number of observations. The result is an x,y coordinate that is at the centroid of all the other points. In allocation with finding the center point of the tornado activity, another pattern that needed analyses  was magnitude.  To do this we used the weighted mean center tool in Arc Map, which does the same operation as a weighted mean center but also allows you to add another factor into the equation, in this case width.  basing the weighted mean center on width pulled the previous geographic mean center of all the tornadoes towards where there were more tornadoes that were larger, and thus more powerful.  These maps can be seen in figure 1 under results

 Another calculation that was made was made using the this data was standard distance operations.  this incorporates both a weighted mean center and adds an circular area that represents the a first order standard deviation of tornadoes occurrences.  What that circle represents is the area where a majority of the tornadoes occurred, and also where the stronger and bigger ones are occurring. These maps can be seen in figure 2 under results


 Results




Figure 2: Three maps which emphasis the mean center and weighted mean center of the Tornadoes in Oklahoma/Kansas. notice the little variance







Figure 2: The emphases of these maps are the standard distances that were applied to each time periods weighted mean center. once again notice the little variance.

County Tornado Statistics

Mean = 4
Standard Deviation = 4.3
range = 0-32



The final analytical procedure employed on this data was to analyze the standard deviation and z scores of the tornado data based on the 2007-20012 county tornado data.  The standard deviation is calculated based off of a single observations variance from the mean. as expected, a majority of the counties fall within the first standard deviation (-.5 - .5).  Similarly, there are less observations that lie outside of the standard deviation.   The calculations presented in this map show a  relatively high number of counties that are above the first standard deviation.

Using both the standard deviation and the mean, students were asked to calculate the z score for three counties.  The z score indicates the actual variance a particular observation deviates from the mean score of all counties.  Using the score of that particular observation you can then find the probability that an observation will occur, with relevance to the data of that time period. The standard deviation of all these counties, as well as the z and p scores of the specified are all illustrated in figure 3.  The percentage associated with the three counties is the probability that a tornado would not occur if the current weather patterns stay the same.

Figure 3: Map showing standard deviation of all counties while also showing the z and p scores of the specified counties.



Conclusion

The overall results that these analytical techniques provided was that between the time periods between 1995-2006 and 2007-2012 the patterns of tornadoes changed very little.  As you can see in the maps of the weighted mean centers and mean centers, the contrasting time periods showed little variation over the period of  17 years.  In regard to the standard deviation, that indicates that a majority of the counties have somewhere between 2-6 tornadoes in their county over a 5 year time period.  As the  map in figure 3 illustrates, there is not much of any pattern that can be seen throughout the area of interest in regards to where more tornadoes are occurring. In a best case scenario I would be able have the count data of tornadoes per county for the 1995-2006 time period for the sake of comparing the change of specific counties. In the 2007-2012 period, only 11% of the counties in the area had 0 tornadoes while the average of each county is 4 tornadoes. 

The implications of these result suggest that tornado patterns haven't changed very much during the the past 17 years of study. The occurrences of these tornadoes are for all intensive purposes, are seemingly very random.  This conclusion came based upon the fact that the mean center is very centrally  located.  The centrality of that point suggest that the geographic occurrences of tornadoes across the area of interest is more or less well dispersed, both in strength and number. In essence what that means is although patterns have not changed,  this model suggests that there is no guarantee of an area being safe from tornadoes.  In regards to all statistics and models calculated, I would advise people to invest in having a tornado shelter.