category: literaturenote citekey: trb08apmlibre title: TRB_08_APM-libre authors: "" zotero_storage: 956ZUCXS collections: imporditud folder: 001_artiklid
By:
Salvatore Cafiso
Associate Professor
Department of Civil and Environmental Engineering
University of Catania
Viale Andrea Doria 6, 95125 Catania, Italy
Phone: +39 0957382213 Fax: +39 0957382247 E-mail: dcafiso@dica.unict.it
Alessandro Di Graziano
Ph.D., P. Eng.
Department of Civil and Environmental Engineering
University of Catania
Viale Andrea Doria 6, 95125 Catania, Italy
Phone: +39 0957382215 Fax: +39 0957382247
E-mail: adigraziano@dica.unict.it
Giacomo Di Silvestro
Ph.D. Candidate, P. Eng.
Department of Civil and Environmental Engineering
University of Catania
Viale Andrea Doria 6, 95125 Catania, Italy
Phone: +39 0957382211 Fax: +39 0957382247 E-mail: gdisilve@dica.unict.it
Grazia La Cava
Ph.D., P. Eng.
Department of Civil and Environmental Engineering
University of Catania
Viale Andrea Doria 6, 95125 Catania, Italy
Phone: +39 0957382223 Fax: +39 0957382247 E-mail: glacava@dica.unict.it
Submission date: 01 August 2007
Words count: 4,721 + 250×10 = 7,221
At European level, road accident fatalities on "two lane rural roads" count up to 56% of the total number of fatalities in road accidents. In Italy, accident data on local rural roads shows a growth of 25 % in the number of accidents from 2001 to 2004.
In this paper a comprehensive procedure starting from low cost data survey to APM implementation, specifically designed for local rural roads, is presented. Firstly, variables significant for two lane local rural highways segmentation in homogeneous sections were identified. Specifically, research efforts aimed to divide roads sample in segments in which all the highway characteristics (exposure, geometrical, consistency and context-related variables) were constant and such to have a minimum length significant for the accident expectation. All the basic data were obtained from surveys purposely set up. Specifically, differential cinematic GPS surveys were used to carry out horizontal alignment and Road Safety Inspections (RSIs) were conducted in order to allow a quantification of the other safety concerns. The homogeneous sections represent the base on which an Accident Prediction Model (APM) for two lane rural road sections was developed. The APMs were defined using Generalized Linear Model approach (GLIM), which has the advantage of overcoming the limitations due to conventional linear regression in accident frequency modeling.
Four APMs were selected with significant goodness of fit. Among these, MODEL 4 even if doesn't represent the best fit of the data, it has the best engineering value as Safety Performance Indicator because significant parameters related to exposure, geometric, consistency and context factors are included in the model.
Keywords: safety indicators, local rural roads, data survey, accident prediction model, low cost
Road Safety represents a priority both to a social level due to the great economic and human costs linked to the consequences of road accidents, and of research for the need of a greater acquaintance and understanding of a phenomenon which for its inner complexity still today introduces wide margins of uncertainty. In this ambit, rural road safety accounts for a considerable share of the total road safety problem.
Statistics showed that deaths on rural roads account for between 28% and 87% of all road fatalities and the majority of the European countries suffer about 60% of their road deaths on rural roads. In particular, fatalities on "no-freeway roads" in rural areas count 56% of the total number of fatalities involved in road accidents (1).
In Italy, the road safety issue is highlighted on local rural roads network and crash data shows a growth of 25 % in the number of accidents and of 21 % in the number of fatalities from 2001 to 2004 (2). This trend goes in opposing to the European Community targets to halve the number of fatalities within 2010 and to the results obtained on the other rural roads typologies (motorways and secondary roads).
Local rural roads are characterized by low-medium traffic flow (usually in the range 1,000 – 6,000 vehicle/day) and short travel distance. The network is composed by two lane roads with design speed in the range 40 - 80 km/h.
Local road network was constructed in the past years using design criteria not suitable with modern safety standards. Moreover, due to the particular complexity of local rural highways on which driver interactions with natural and anthropic context are higher and to the problem of budget limits of Local Road Agencies, restoration interventions planning on these roads, hardly can be linked to the design standards check which not ever are suitable in economic terms for benefits- costs analysis. For this reason, on these roads, safety improvements must be related to a comprehensive approach taking into account the safety performance of the road. Accident Prediction Models (APMs) represents a direct method to analyze how safety is related to different road features and characteristics.
The need to achieve and collect highway features data significant to implement and to use an APM on this type of roads is emerging as a considerable problem. Aside, only few studies (3, 4, 5) have attempted to highlight the value to purposely define homogeneous sections in roads sample on which accident modeling are developed. Furthermore these studies didn't have awareness to the definition and characterization of the procedure for selecting homogeneous sections.
In this paper a comprehensive procedure starting from low cost data survey to APM implementation, specifically designed for local rural roads, is presented. Firstly, variables significant for two lane local rural highways segmentation in homogeneous sections were identified. Specifically, research efforts aimed to divide roads sample in segments in which all the highway characteristics (exposure, geometrical, consistency and context-related variables) were constant and such to have a minimum length significant for the accident expectation. A GPS survey was used to carry out horizontal alignment and cross sectional information and Road Safety Inspections (RSIs) (6, 7,8,9) were conducted in order to allow a quantification of the other safety concerns.
Four APMs were defined using Generalized Linear Model approach (GLIM), which has the advantage of overcoming the limitations due to conventional linear regression in accident frequency modeling. Each APM uses a different set of explanatory variables and was defined with different criteria (homogeneous section check, best goodness of fit, best significant variables and best engineering value).
For the safety performance evaluation the following basic highway features were identified:
Even if other useful information could be used (sight distance, gradient, pavement condition, etc…) the previous listed parameters are particularly relevant and can be relatively easy collected. With respect to Italian situation, due to the partial or the total lack of the as-built drawings of rural local roads, an alternative, cost effective and practical method for collecting road data was required.
A GPS survey was used to carry out horizontal alignment information (item 1, 2 and 3) and Road Safety Inspections (RSIs) (6, 7, 8, 9) were conducted in order to allow a quantification of the other safety concerns (item 4, 5 and 6). The GPS investigation is performed travelling by car the lane axis at moderate speed (60 km/h) and using a L2 GPS equipment in cinematic-differential method (DGPS), so as to obtain an error of less than 10 cm. The GPS survey allows a series of points to be identified along the trajectory of the vehicle (lane axis). Although this information makes it possible to produce an efficient representation of the stretch, it is not enough for the parameters of the alignment design to be identified. Therefore, the next phase of the procedure consists in elaborating the data so as to identify the single element of the axis (tangents, circular curves, clotoids) and determine their geometric parameters (length, radius, angular extension). The set up model uses regression splines with the aid of suitable smoothing factors to correct the inevitable axis survey errors due to the actual trajectory followed by the vehicle or to measurement errors (10). In particular, smoothing functions are used on the spline to correct the curvature locally by means of a P parameter whose value varies between 0 (the spline represents a minimum squared linear regression) and 1 (the spline represents an interpolation spline passing exactly through the points). Operationally, a parameter P < 1 is defined limiting the smoothing error to 1 m, or rather, the spline carried out has a maximum of 1 m distance from the surveyed GPS points. The coefficients of the cubic spline were defined as
$$f_i(x)=a_i x^3 + b_i x^2 + c_i x + d_i,$$ (1)
where
ai, bi, ci, di = spline coefficients
the curvilinear extension (s) of the spline from the point of origin of co-ordinate (x0, ƒ(x0)) to the point of the co-ordinate (x, ƒ(x)) is calculated according to the following equation:
$$s = \sum{i} \int{x0}^{x} \sqrt{1 + f{i}'(x)} dx \quad [m]$$ (2)
where
ƒi ′(x) = first derivative of the spline function ƒi(x)
The curvature (1/ρ) can be calculated by means of the following equation:
$$\frac{1}{\rho} = \frac{f''(x)}{\left(1 + f'(x)^2\right)^{3/2}} \qquad [1/m] \tag{3}$$
where
ƒi"(x) = second derivative of the spline function ƒi(x).
When the curvature function (s, 1/ρ) is carried out, a definition of the actual geometrical elements of the alignment is obtained by looking for the succession of defined geometric elements (straight stretches, clotoids, circumferences) that guarantee the same angular and linear extension of the spline curvature. From this condition, an equality between the area under the spline curvature (curvilinear shape) and the area under the curvature of road alignment (trapezium shape), is imposed. Therefore, the numerical values of the radius R of the curve and the parameter A of the clotoid can be calculated minimizing the root mean square sum (RMS) of the error between the curvature diagrams of spline and alignment (see Figure 1).
FIGURE 1 Comparison between the spline and the reconstructed curvature.
For old roads constructed without the use of clotoids, like those in this research, constant curve elements (rectangular shape) can be identified.
Roadside hazard and driveway density evaluation was carried out using the methodologies proper of the RSIs. Considering that at present different procedures are adopted both in some EU countries and no EU ones, with the aim to improve the effectiveness and the reliability of the methodology of RSIs, a new operating inspection procedure was identified and formalized as part of the IASP project carried out by the University of Catania and co-funded by the Regional Province of Catania and by the European Commission (General Directorate Energy and Transport) (11).
According to specific and detailed criteria reported in the Inspection manual (8), a team of safety experts expressed judgment on various safety issues using checklists that were filled in both directions with a step of 200 m.
Safety issues are ranked as 'high level problem' (score equal to 2), 'low level problem' (score equal to 1) and 'no problem' (score equal to 0) (6, 9). In IASP procedure, a part of the checklists that were compiled during the inspection is devoted to roadside with particular reference to specific factors, as shown in Table 1 and another part is related to driveway presence along the road under inspection. Lane and shoulder width are qualitatively identified as "high, low, no problem" and direct measures are surveyed with respect to homogeneous sections.
0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 1.8 2.0 Roadside Embankments Bridges Dangerous terminal and transitions Trees, utility poles and rigid obstacles Ditches
TABLE 1 Part of the Checklist for devoted to Roadside Evaluation
For the present research, a sample of about 92.280 km of two lane local rural roads located in the province of Catania and Ragusa (Italy) was used in order to check the proposed procedure and to perform the APM.
Usually, a time interval between two and five years can be considered as reference accident history. Longer temporal interval has to be avoided because they could produce the refusal of the hypothesis of Poisson distribution for the accident frequency dues to a non stationary phenomenon. A Five years of analysis was chosen as accident period to compensate the low traffic flow usually expected on local rural roads.
Road accident records, in which at least one person died and/or was injured, were acquired from a survey directly carried out from Road Traffic Police, Municipal Police and Carabinieri reports. All the accidents were located to the geometric elements (e.g. curve or tangent) along the road using mile stones references listed in the accident report. Accident records located in built up areas and at intersections were not used for the APM development. A total of 150 accidents were reported for the period 1999-2003, with a total of 293 injured and 9 dead people.
Starting from data collected on the road, some new parameters were defined in order to describe those characteristics considered fundamental to segment the road into homogeneous sections:
In order to single out the road sections with homogenous horizontal alignment characteristics reference was made to the indications present in the German standard (12). This provides a method for singling out homogenous road sections by evaluating the Cumulative Curvature Change Rate
$$CCR = \Sigma_i |\gamma_i| ) / L \qquad [gon/km]$$ (4)
The sum of the deflection angles (γi) of contiguous elements "i" of the horizontal alignment is represented in a diagram as a function of distance (L) (Figure 2).
It is quite easy to identify the sections in which to subdivide the road. They are characterized by various constant slopes interpolating the overall curvature. Based on the German procedure the length of the homogenous sections has to be always higher than 2 km.
FIGURE 2 Cumulative change rate.
The Average Paved Section (W) was computed reporting a new measurement each time a change of 50 cm was surveyed.
The AADT was analysed in terms of network, assigning a constant value to each segment delimited by two main intersection with a significant flow variability.
Referring to the Road Side Hazard (RSH), as it was evaluated during RSIs with a step unit of 200 m by the checklists filled in both directions (right and left side) with reference to different roadside items, first of all, a data aggregation was necessary.
A weighted score of RSH for each inspection step was defined as follows:
$$RSH{ik} = \frac{\sum{k=1}^{2} \max{i} (Score{ik} \times Weight_{i})}{2}$$ (5)
where:
k = direction of the survey (1 = right side; 2 = left side)
Scoreik = score of the roadside safety items i in the inspection units (0, 1, or 2);
Weighti = relative weight of the roadside safety item i (see Table 2).
TABLE 2 Relative Weights of the Roadside Safety Items
| Detailed Safety Issue | Relative Weight |
|---|---|
| Embankments | 3 |
| Bridges | 5 |
| Dangerous terminals and transitions | 2 |
| Trees, utility poles and rigid obstacles | 2 |
| Ditches | 1 |
In order to determine a sequence of RSH values that can be connected, with statistical meaning, to identify an homogeneous road section, RSH data were treated defining reference constant values that permit a minimization sum of squared errors (13). The method allows to define among all feasible set of segmentations that one which minimize the sum of squared error in function of a minimum length Lmin. Then a following aggregation is carried out taking into account contiguous segments with equal mean in accordance with the t-student test. In this application a minimum length Lmin equal to 1000 meters and a t-test significance level of 85% were selected (Figure 3)
FIGURE 3 RSH Segmentation.
Next to the identification of parameters characterising the road, a segmentation in homogenous sections was carried out by any of the four descriptive average parameters (CCR, W, AADT, RSH). The final homogeneous road section is a section where all the average parameters are constant. In Figure 4 an example of a road of the investigated sample, with a length of 9 km, subdivided in 8 homogeneous sections is shown.
FIGURE 4 Road homogeneous sections.
Referring to the sample under study 62 homogeneous sections on a total length of 92.280 km of roads were defined. The length of these sections ranges from a minimum of 151 m to a maximum of 4141 m.
Considering that the final target of this research is the implementation of an APM based on actual accidents, the procedure of road segmentation can produce homogeneous sections that are not significant of length and traffic flow with respect to accident exposure. The level of exposure (million vehicles per kilometres) in the investigation period (five years) could be too low to produce a significant expected number of accidents. Therefore, it can be assumed, with respect to the exposure, that at least one accident must be expected in each section. Therefore, in this first phase of the research, to check the consistency of the road segmentation a basic model with only the exposure independent variables was carried out using generalized linear model (GLIM):
Iteratively some sections were discarded and then other contiguous ones were aggregated in order to guarantee at least a minimum of one predicted accident (E(Y) > 1) in each road sections. In this way it was possible to obtain sections with minimum length of 500 m (AADT= 6,700 vehic./day) and a maximum one of 4141 m (AADT= 600 vehic./day).
The base model used to check the minimum section length is the following:
$$MODEL 1: E(Y) = e^{-14.14} \times AADT^{0.985} \times L$$ (6)
where:
E(Y) = Predicted accident frequency / 5 years of the random variable Y;
L = length of the segment under consideration (km);
AADT = Average Annual Daily Traffic (veh/day);
Starting from the 62 homogeneous sections resulting from the road segmentation, the procedure produced 55 segments for the development of APMs as explained in the following paragraphs.
Defined the basic homogeneous sections, other variables were, afterwards, associated to the road sections in order to perform the APM. These variables could be significant to explain the APM but their variability does not influence the homogeneity of the section because they derived from the basic descriptive parameters used to define the basic segmentation. These variables could be related to the geometric and operational condition (CR, TR, Vavg), to the design consistency (Ra, σ, ∆Vn, ∆V10, ∆V20) and to the context (DD) of the road as following explained.
Referring to geometric and operational variables Curvature Ratio (CR) and Tangent Ratio (TR) were computed by the way of following equations:
$$\mathbf{CR} = \frac{\sum{j=1}^{k} L{Cj}}{\mathbf{L}_{HS}} \tag{7}$$
$$TR = \frac{MAX{e=1}^{t}(L{T})}{L_{HS}}$$ (8)
where:
LHS = total length of the homogeneous section [km]
LCj = length of jth curve in the homogeneous section [km];
LTe = length of eth tangent in the homogeneous section [km]
The Average Operating Speed (Vavg) was computed on the base of the profile of operating speed as the average weighted speed along the entire segment of length L. Operating speed on tangents and curves were developed using an prediction model for local rural roads in Italy (9). Then, in order to define a realistic operating speed profile, it was necessary to assume that the acceleration and deceleration phases between tangent and curve lasted for 2 seconds on the tangent and 1second on the curve.
Referring to design consistency variables, the two measures of Relative Area (Ra) bounded by the speed profile and the average weighted speed and Standard Deviation (σ) of operating speeds profile were calculated (14) by the way of the following equations:
$$R{a} = \frac{\sum{i=1}^{n} a{i}}{L{HS}} [m/s]$$ (9)
$$\sigma = \sqrt{\frac{\sum_{i=1}^{n} (Vi - V{avg})}{n}} \text{ [km/h]}$$ (10)
where:
ai = bounded area between the speed profile and the average operating speed line [m2 /s];
Vi = operating speed of the ith geometric element [km/h];
Vavg = average weighted speed along the entire segment of length L [km/h];
n = number of geometric elements along a section.
Other three consistency variables were selected referring to the speed differential (∆Vs) between contiguous elements in the homogeneous section, by the way of the following equations:
$$\Delta V{n} = \frac{\sum{s=1}^{n{\Delta V}} \Delta V{s}}{n_{\Delta V}} [km/h]$$ (11)
$$\Delta V{10} = \frac{N(\Delta V > 10)}{L{HS}} [1/km]$$ (12)
$$\Delta V{20} = \frac{N(\Delta V > 20)}{L{HS}} [1/km]$$ (13)
where:
n∆V = number of speed differential (∆Vs) in the homogeneous section (HS);
N(∆V>10) = number of speed differential (∆Vs) higher than 10 km/h in the HS;
N(∆V>20) = number of speed differential (∆Vs) higher than 20 km/h in the HS.
Figure 5 shows an example of operating speed profile and variables measures related to horizontal alignment.
FIGURE 5 Operating speed profile and connected variables.
Referring to the context-related variables, as direct accesses to roads show the dramatic effect on safety and can significantly increase accidents (15), the driveway density (DD) was considered. The DD was obtained from the checklists as previous described.
In the past, several APMs have been developed using conventional linear regression, which assumes a normal error structure for the response variable, a constant variance for the residuals, and the existence of a linear relationship between the response and explanatory variables. However, several researchers (16, 17, 18, 19, 20) have demonstrated the inappropriateness of conventional linear regression for modeling discrete, nonnegative, and rare events such as traffic accidents. These researchers have also recommended the assumption of a Poisson or a Negative Binomial error structure when modeling road accidents.
The models proposed by the authors are based on the Generalized Linear Model approach (GLIM), which has the advantage of overcoming the limitations due to conventional linear regression in accident frequency modelling. The GLIM approach used in this study, is based on above mentioned literature references and the regression analyses were performed by use of the software package GenStat 7.2.
The general form of the accident prediction model adopted is:
$$E(Y) = e^{a_0} \cdot L \cdot AADT^{a1} \cdot e^{\sum{j=1}^{m} b_j \cdot x_j}$$ (14)
where
E(Y) = Expected accident frequency /5 years of the random variable Y;
L = length of the segment under consideration (km); AADT = Average Annual Daily Traffic (AADT) (veh/day);
xjAny of m additional variables;
a0, a1, and bj = model parameters.
Several measurements are usually used to assess the goodness of fit of the model and the significance of the model parameters.The reported indicators are the t-ratio for the model parameters, the ф value of the dispersion parameter, the scaled deviance (SD), the Pearson χ statistics and the χ 2 0.05, dof distribution for the given degrees of freedom and level of confidence (see Table 6).
The SD is defined as the likelihood ratio test statistic measuring twice the difference between the log likelihoods of the studied model and the full or saturated model. The full model has as many parameters as there are observations so that the model fits the data perfectly. Therefore, the full model, which possesses the maximum log likelihood achievable under the given data, provides a baseline for assessing the goodness of fit of an intermediate model with parameters. McCullagh and Nelder (21) have shown that for negative binomial error structure the scaled deviance is as follows:
$$SD = 2\sum_{i=1}^{n} \left[ yi \ln\left(\frac{yi}{\hat{E}(yi)}\right) - (yi+k)\ln\left(\frac{yi+k}{\hat{E}(yi)+k}\right) \right]$$ (15)
where:
SD = scaled deviance;
yi = observed number of accidents in the segment i;
*E*ˆ (yi) = predicted number of accidents in the segment i;
k = the negative binomial parameter.
For a well-fitted model, both the scaled deviance and the Pearson χ should be significantly compared with the critical value obtained from the χ 2 distribution for the given degrees of freedom and level of confidence.
The Pearson χ 2can be calculated by means of the following formula:
Pearson $$\chi^2 = \sum_{i=1}^n \frac{\left[yi - \hat{E}(yi)\right]^2}{Var(yi)}$$ (16)
where:
Var(yi) = variance of the observed accidents.
In order to verify the goodness of fit of the model, these two measures must be compared with the value obtained from χ 2distribution with the sample degrees of freedom. For the model to be considered significant, both the two measures should be less than χ critical value.
Another measure of model goodness of fit is the dispersion parameter Φ :
$$\Phi = \frac{\text{Pearson}\chi^2}{\text{n-p}} \tag{17}$$
where:
n = Number of observations;
p number of model parameters to be estimated.
As above shown, Φ can be obtained by dividing the Pearson χ by n-p, and it is a useful measure for assessing the fit of model. A value near 1.0 means that the Poisson error assumption of the model is equivalent to that found in the observed data.
All the model were firstly defined using a Poisson error structure (POISSON), when dispersion parameter Φ resulted much greater than 1, a model refinement was performed adopting a negative binomial error distribution (NB)
A summary statistics of all the variables deriving from data treatment and elaboration and used for the APM implementation are reported in Table 3.
TABLE 3 Variables Groups for the APM Implementation
| Variables groups | Abbreviation | Mean | Min | Max | Standard deviation |
|---|---|---|---|---|---|
| EXPOSURE | |||||
| LHS | 1658 | 500 | 4141 | 829 | |
| AADT | 2708 | 600 | 6700 | 1450 |
| GEOMETRIC AND OPERATIONAL | | | | | | | | | | |---------------------------|------|--------|------|--------|--------|--|--|--|--| | | CCR | 128.02 | 2.20 | 629.21 | 152.17 | | | | | | | W | 6.99 | 6.10 | 8.50 | 0.87 | | | | | | | TR | 0.37 | 0.04 | 1.00 | 0.28 | | | | | | | CR | 296.49 | 0.00 | 658.23 | 172.30 | | | | | | | Vavg | 92.72 | 57.5 | 115.5 | 17.1 | | | | | | CONSISTENCY | | | | | | | | | | | | Ra | 1.29 | 0.00 | 3.77 | 0.99 | | | | | | | σ | 7.21 | 0.00 | 16.45 | 50.03 | | | | | | | ∆Vn | 9.77 | 0.00 | 24.32 | 6.08 | | | | | | | ∆V10 | 3.72 | 0.00 | 14.30 | 4.30 | | | | | | | ∆V20 | 1.75 | 0.00 | 12.55 | 2.82 | | | | | | CONTEXT-RELATED | | | | | | | | | | | | RSH | 1.57 | 0.00 | 4.93 | 1.04 | | | | | | | DD | 8.19 | 2.22 | 21.81 | 4.68 | | | | |
The purpose of the analysis is to identify an explanatory model to characterize the relationship of each predictor to the expected number of accident (outcome variable). To achieve this goal, the identities of the variables in the model are critical, and the analyst must take great care in choosing which variables under study significantly affect accident frequency. For this reason a preliminary analysis was carried out to identify the correlation between two variables among those reported in shown in Table 4.
The most common measure of correlation is the Pearson Product Moment Correlation (Pearson's correlation). Pearson's correlation reflects the degree of linear relationship between two variables. It ranges from +1 to -1. A correlation of +1 means that there is a perfect positive linear relationship between variables.
Numbers in parenthesis in each location of the table 4 are P-value which tests the statistical significance of the estimated correlations. P-values below 0.1 indicate statistically significant non-zero correlations at the 90% confidence level.
| Variables | LHS | AADT | RSH | DD | CCR | W | Vavg | s | Ra | DV10 | DV20 | DVn | CR | TR | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CP | |||||||||||||||
| LHS | Sig. | 1 | |||||||||||||
| CP | -0,1240 | ||||||||||||||
| AADT | Sig. | (0.3670) | 1 | ||||||||||||
| RSH | CP | 0,0050 | -0,0741 | 1 | |||||||||||
| Sig. | (0.9711) (0.5906) | ||||||||||||||
| DD | CP | -0,0321 | 0,1327 | -0,0497 | 1 | ||||||||||
| Sig. | (0.8153) (0.3341) (0.7184) | ||||||||||||||
| CCR | CP | 0,1061 | -0,3528 | 0,1948 | -0,0997 | 1 | |||||||||
| Sig. | (0.4408) (0.0082) (0.1539) (0.4687) | ||||||||||||||
| W | CP | -0,0814 | 0,0205 | 0,1684 | 0,1770 | -0,3954 | 1 | ||||||||
| Sig. | (0.5547) (0.8821) (0.2190) (0.1961) (0.0028) | ||||||||||||||
| Vavg | CP | -0,1526 | 0,3738 | -0,1616 | -0,0189 | -0,8850 | 0,5364 | 1 | |||||||
| Sig. | (0.2659) (0.0049) (0.2382) (0.8916) (0.0000) (0.0000) | ||||||||||||||
| s | CP | 0,2580 | -0,1614 | 0,2010 | 0,1737 | 0,7773 | -0,4531 | -0,8846 | 1 | ||||||
| Sig. | (0.0573) (0.2392) (0.1410) (0.2050) (0.0000) (0.0005) (0.0000) | ||||||||||||||
| Ra | CP | 0,2012 | -0,3112 | 0,2177 | 0,0752 | 0,8722 | -0,4396 | -0,9589 | 0,9522 | 1 | |||||
| Sig. | (0.1403) (0.0207) (0.1106) (0.5837) (0.0000) (0.0008) (0.0000) (0.0000) | ||||||||||||||
| DV10 | CP | 0,1089 | -0,2442 | 0,1859 | -0,0015 | 0,8690 | -0,5320 | -0,9132 | 0,8398 | 0,8944 | 1 | ||||
| Sig. | (0.4294) (0.0725) (0.1742) (0.9913) (0.0000) (0.0000) (0.0000) (0.0000) (0.0000) | ||||||||||||||
| DV20 | CP Sig. |
0,0775 | -0,2303 | 0,2388 | -0,0750 | 0,8463 | -0,4228 | -0,8179 | 0,7692 | 0,8351 (0.5743) (0.0909) (0.0789) (0.5851) (0.0000) (0.0013) (0.0000) (0.0000) (0.0000) (0.0000) |
0,8916 | 1 | |||
| CP | 0,2264 | -0,1888 | 0,2536 | 0,1017 | 0,7728 | -0,4344 | -0,8483 | 0,9262 | 0,9141 | 0,8305 | 0,8357 | ||||
| DVn | Sig. | (0.0964) (0.1673) (0.0617) (0.4601) (0.0000) (0.0009) (0.0000) (0.0000) (0.0000) (0.0000) (0.0000) | 1 | ||||||||||||
| CP | 0,0413 | -0,5961 | 0,0925 | 0,0284 | 0,7259 | -0,2624 | -0,8011 | 0,6068 | 0,7375 | 0,6707 | 0,5672 | 0,5521 | |||
| CR | Sig. | (0.7647) (0.0000) (0.5015) (0.8374) (0.0000) (0.0529) (0.0000) (0.0000) (0.0000) (0.0000) (0.0000) (0.0000) | 1 | ||||||||||||
| CP | -0,2417 | 0,4737 | -0,0880 | -0,0717 | -0,6006 | 0,4590 | 0,7986 | -0,6827 | -0,7509 | -0,6431 | -0,5098 | -0,6476 | -0,8371 | ||
| TR | Sig. | (0.0760) (0.0003) (0.5157) (0.6114) (0.0000) (0.0004) (0.0000) (0.0000) (0.0000) (0.0000) (0.0001) (0.0000) (0.0000) | 1 |
TABLE 4 Pearson correlation coefficients between variables.
Table 4 gives an idea about which variables were correlated (CP close to +1 and P-value below 0.1). In this way it can be possible to expect that not correlated variables can be significant for the development of the APM. These couple of variables are identified in table 4.
The number of variables in the model was reduced by using a backward selection procedure. At the beginning all the variables were entered into the model, then they were sequentially discharged starting with the variable having the weakest association with the outcome and continuing until the only variables left in the model were those related to the outcome variableat a significance level of 85% (i.e. P< 0.15). The P-values of the t statistics test the hypothesis that a regression coefficient is 0 when the other predictors are in the model.
The accident prediction model (MODEL 2) obtained was:
MODEL 2: $$E(Y) = e^{-16.86} \cdot L{HS} \cdot AADT^{1.074} \cdot e^{0.190 \cdot RSH + 0.071 \cdot DD + 0.124 \cdot \Delta V{10} + 0.206 \cdot W - 0.114 \cdot \sigma}$$ (18)
This model highlights the influence besides the exposure variable also of the context-related variables RSH and DD, consistency variables DV10 and σ, and among the geometric variables only B was included in the model. However, it can be noted that the model parameter of W and σ are respectively positive and negative. This means that increasing the paved section (W) and decreasing the speed standard deviation (σ) an accident growth is expected.
This result even if significant from a statistical point of view is not logical from an engineering judgment and it can be explained considered that W and σ can be correlated with other road features.
Table 4 shows that wider road paved sections (W) and lower speed deviances (σ) are associated to higher operating speed that can justify the increase in accident prediction.
If both W and σ are removed from the model the backward process generates a new model (MODEL 3) with a very strong significance of the description parameters L, AADT, RSH and DD.
MODEL 3: $$E(Y) = e^{-14,490} \cdot L_{HS} \cdot AADT^{0,943} \cdot e^{0,188 \cdot RSH + 0,043 \cdot DD}$$ (19)
Finally, starting from MODEL 3 a forward process was adopted to include in the model other variables that can be indicators of the road safety performance:
MODEL 4: $$E(Y) = e^{-16,720} \cdot L{HS} \cdot AADT^{0,859} \cdot e^{0,202 \cdot RSH + 0,053 \cdot DD + 0,089 \cdot \Delta V{10} + 0,026 \cdot V_{avg}}$$ (20)
Even if MODEL 4 doesn't represent the best fit of the data, it has a significant goodness of fit and it produces significant parameters related to exposure, geometric, consistency and context factors.
The model parameters are listed in Table 5 with regression coefficients and the model goodness of fit statistics.
| MODEL
ID | DISTR. | | Df VARIABLES | PAR. | ESTIMATE | t-ratio | t-pr | DISPERSION
PARAMETER | SD | 2
Pearsonc | 2
c
0,05(GdL) | |
|-------------|------------|----|--------------|------|----------|---------|----------------|-------------------------|--------|---------------|---------------------|--|
| | | | COST | a0 | -14.14 | | -9.170 < 0.001 | | | | | |
| 1 | NB | 52 | LHS | a1 | 1 | | | 1.250 | 54.552 | 66.250 | 70.993 | |
| | | | AADT | a2 | 0.985 | 5.100 | < 0.001 | | | | | |
| | | | COST | a0 | -16.86 | | -8.530 < 0.001 | | | 52.320 | 65.171 | |
| | | | LHS | a1 | 1 | | | | | | | |
| | | | AADT | a2 | 1.074 | 5.770 | < 0.001 | | | | | |
| 2 | POISSON 48 | | RSH | b1 | 0.190 | 2.080 | 0.042 | 1.090 | | | | |
| | | | DD | b2 | 0.071 | 3.390 | 0.001 | | 48.037 | | | |
| | | | DV10 | b3 | 0.124 | 3.250 | 0.002 | | | | | |
| | | | W | b4 | 0.206 | 1.560 | 0.124 | | | | | |
| | | | s | b5 | -0.114 | -3.330 | 0.002 | | | | | |
| | | | COST | a0 | -14.490 | | -9.840 < 0.001 | | 53.894 | 57.630 | 68.669 | |
| | NB | | LHS | a1 | 1.000 | | | 1.130 | | | | |
| 3 | | 51 | AADT | a2 | 0.943 | 5.120 | < 0.001 | | | | | |
| | | | RSH | b1 | 0.188 | 1.900 | 0.062 | | | | | |
| | | | DD | b2 | 0.043 | 2.110 | 0.040 | | | | | |
| | | | COST | a0 | -16.720 | | -8.080 < 0.001 | | | | | |
| | NB | | | LHS | a1 | 1.000 | | | | | | |
| | | | AADT | a2 | 0.859 | 4.260 | < 0.001 | 1.110 | 48.829 | 54.390 | 66.339 | |
| 4 | | 49 | RSH | b1 | 0.202 | 1.990 | 0.052 | | | | | |
| | | | DD | b2 | 0.053 | 2.440 | 0.018 | | | | | |
| | | | DV10 | b3 | 0.089 | 1.460 | 0.150 | | | | | |
| | | | Vavg | b6 | 0.026 | 1.730 | 0.091 | | | | | |
TABLE 5 Model Parameters and Indicators for Model Goodness of Fit
A comprehensive procedure including low cost data survey, road segmentation methodology and APM development, specifically designed for local rural roads, is presented. Due to the low traffic volume and to the lack of significant number of accidents, the segmentation of road stretches into homogeneous sections is required to develop an Accident Prediction Model (APM), specifically designed as Safety performance indicator of local rural roads. All the road data were obtained from differential cinematic GPS surveys and Road Safety Inspections (RSIs).
Curvature Change Rate, average roadway Width (W), Average Annual Daily Traffic, Road Side Hazard rating were identify as significant variables for the highways segmentation in homogeneous sections. The segmentation procedure aimed to divide road sample in segments in which all the highway characteristics (exposure, geometrical, consistency and context-related variables) were constant and such to have a minimum length significant for the accident expectation.
Considering that the final target of this research was the implementation of the APM, accident data were collected and a total of 14 variables subdivided in four groups (exposure, geometric, consistency and context) were identified and related to the homogeneous sections previously defined.
The APM was carried out using Generalized Linear Model approach (GLIM), which has the advantage of overcoming the limitations due to conventional linear regression in accident frequency modeling.
Four APM models were selected with significant goodness of fit.
MODEL 1, using as explanatory variables AADT and L, is a basic model to check the minimum length of t homogeneous section.
MODEL 2, was carried out with a backward procedure selecting as explanatory variables AADT, L, RHS, DD, W, DV10 and s.
When both W and s were removed from MODEL 2, the backward process generates a new model (MODEL 3) with a very strong significance of the description parameters L, AADT, RSH and DD.
Although MODEL 4 doesn't represent the best fit of the data, it has the best engineering value as Safety Performance indicator because significant parameters related to exposure, geometric, consistency and context factors are included in the model.
The authors would like to acknowledge the Road Agencies of the Province of Catania and Ragusa for supporting this research.