Los Angeles is a city of around 4 Mio1 inhabitants from diverse backgrounds (about half2 of the population is latino) and had a crime rate of about 6353 per 100’000 inhabitants per year in 2015. The city hosts the biggest port complex in the US and is an illegal drugs hub4 for the country. Finding a way to foster this diversity so as to take advantage of the city’s economical potential and reduce crime would provide more room for children to develop, achieve high education levels and freely roam around the city.

The aim of this paper is to identify how demographic data at the zip code level in Los Angeles impact crime rates in 2010. More precisely, thw two following hypotheses are formulated and tested in the following:

For the purpose of the study, three data sets are used:

In the following, non-spatial regression analysis is used. This provides the basis for further exploration with spatial methods, i.e. bivariate spatial maps, centrographic statistics and LISA statistics based on Moran’s I. First, an exploratory data analysis is performed in order to assess the distribution of the variables of interest.

Descriptive Statistics

The histograms below (see Figure @ref(fig:histograms)) show that average household size and median age are centered around 3 and 40 respectively. Crime rate has a high number of 1, 2 and 3 values (the underlying data has no 0 values). This might be due to the lack of data in certain regions. Those 1s values were not dismissed in order to not dismiss zip code areas arbitrarily in the final analysis, instead the spatial plots use quantiles, which allow to categorize all low values together and allow for comparison with higher and more reasonable estimates of crime rates.

The interaction between the explanatory variables (median age and average household size) and the dependent variable (number of crimes per zip code) can be seen on the plot below. The natural log of number of crimes was used in order to standardize the data, since it is originally left skewed. It seems that that no particular relationship exists on a first look, this will be tested by the regressions below.

non-spatial bivariate regressions, non-spatial correlations

Below are two regression analyses for the two independent variables (average household size and median age). It is deceiving in terms of explained variance and of the explanatory value of the independent variable median age. Average household size is significantly different from zero at a 5% confidence level though: a bigger household would imply more crimes as stated in the hypothesis in the introduction.

Dependent variable:
log(crimes)
Average.Household.Size 0.675**
(0.305)
Median.Age 0.005
(0.028)
Constant 3.398**
(1.441)
Observations 242
R2 0.020
Adjusted R2 0.012
Residual Std. Error 3.993 (df = 239)
F Statistic 2.457* (df = 2; 239)
Note: p<0.1; p<0.05; p<0.01

Those results are not surprising given the scatter plot shown in the previous section. This relative lack of relationship especially for median age does not entail that no trends can be found by treating the data as spatial data. This will be investigated further below.

2 variable maps

Number of Crimes VS Average Household Size

The spatial data delivers more detailed information on certain underlying relationships. Under the hypothesis stated in the introduction, one would expect that the more a circle tends to reddish (i.e. the higher the average household size), the more an area tends to blueish (the higher the number of crimes). This is however not the case, except for three areas north-west of Santa Monica.