To make high-quality research more accessible and easier to explore.

Fields:
2 results ✕ Clear filters

Multicollinearity: A Bayesian Interpretation

The Review of Economics and Statistics 1973 55(3), 371
T HE problem of collinear data sets in the context of the univariate multiple regression model has generated a confusing array of papers, comments and footnotes. Perusal of this literature does not lead the reader to conclude the problem has even been rigorously defined. The purpose of this paper is to offer several rigorous definitions and to suggest several substantive quantitative summaries of the degree of the problem. Specifically we address our attention to the regression model Y = X,B + u where ,B is a k-dimensional parameter vector with elements 81,12, ... *k, where Y and u are T X 1 vectors and where X is a T X k matrix. Inferences are to be made about the vector ,B from observations of Y and X. When the columns of X are orthogonal, the design matrix X'X is diagonal. Correlated columns of X imply a nondiagonal design matrix. The collinearity problem has to do with the differences in the inferences may be drawn in these two situations. The principal claim of this paper is the most important aspects of the collinearity problem derive from the existence of undominated prior information which causes major problems in interpreting the data evidence. It is claimed here if our a priori knowledge of parameter values were either certain or completely uncertain the aspects of the collinearity problem most of us worry about would disappear.' As an empirical test of this proposition consider the situations when collinearity is identified as a culprit. Usually signs are wrong or point estimates are otherwise peculiar. Occasionally confidence intervals overlap unlikely regions of the parameter space. Yet to say these things is to say there exists undominated prior information. Classical inference, with the possible exception of the pretesting literature, necessarily excludes undominated prior information. As a result most discussions of the collinearity problem miss a critical point. The textbook discussions including Theil (1971, p. 149), Malinvaud (1970, p. 218), and Goldberger (1964, p. 192), observe when the design matrix X'X becomes singular, the least squares estimator is non-unique and the sampling distribution has finite variance only for certain estimable functions. Thus extreme collinearity is implicitly defined as total lack of sample information about some parameters. The case of less extreme collinearity is not dealt with so trivially since there is nothing in the least-squares theorems is obviously dependent on the near non-invertibility of the design matrix. This fact has led Kmenta (1971, p. 391) to conclude that a high degree of multicollinearity is simply a feature of the sample contributes to the unreliability of the estimated coefficients, but has no relevance for the conclusions drawn as a result of this unreliability. To put this another way, the problem of defining collinearity may be solved by identifying a distance function for measuring the closeness of the design matrix to some noninvertible matrix in which the collinearity problem is unambiguously extreme. Since the extreme case is associated with infinite marginal variances on the parameters, authors such as Theil (1971, p. 152), Malinvaud (1970, p. 218), and Goldberger (1964, p. 193) use a distance function informally related to the sampling variance of the coefficients. Collinearity is defined as large variances. The failure of this definition is instead of defining a new problem, it identifies a new cause of an already well-understood problem weak evidence. Although collinearity as a cause of the weak evidence problem can be distinguished from other causes such as small samples or large residual error variances, collinearity as Received for publication November 7, 1972. Revision accepted for publication January 23, 1973. * Research for this paper was supported by NSF Grant GS 319.29. The author has benefited from conversations on the subject with Gary Chamberlain and Richard Kopcke. Both the discussion and the content have benefited significantly from a referee's comments. An earlier version was presented at the NBER-NSF Seminar in Bayesian Inference in Econometrics at the University of Minnesota, October, 197,2. 1 See Zellner (1971, chapter 2) for a discussion of the problem of defining complete uncertainty.

Empirically Weighted Indexes for Import Demand Functions

The Review of Economics and Statistics 1973 55(4), 441
T HIS paper presents an empirical approach to the aggregation problem. An aggregate dependent variable is properly thought to depend on a long string of disaggregated explanatory variables. There are two extreme ways of analyzing an equation such as this. We may use all of the explanatory variables individually, or we may summarize them in an index and use only the index as an explanatory variable.1 Both of these approaches have serious shortcomings and the purpose of this paper is to consider methods that lie between these extremes. The difficulties with the two extreme approaches are well known. On the one hand data limitations in the form of degrees of freedom inadequacies or multicollinearity usually completely rule out the purely empirical approach. On the other hand, the use of indexes as explanatory variables implies a specification error with unhappy implications worked out by Theil (1954 and 1971), Klein (1946), Allen (1965) and others. It is useful to consider the role of prior judgement in these two approaches. The one analysis excludes prior information about the coefficients and estimates would depend entirely on data evidence if that were possible. The other approach is based on the assumption that the coefficients are known with certainty (up to a factor of proportionality) and no data evidence can alter that constraint. In this paper we will pursue an intermediate course that involves weakening the assumption that the index weights are known with certainty but not to the point that the weights are totally unknown. More formally, this means formulating a proper prior distribution on the coefficients in the unconstrained equation and updating that prior via Bayes Rule. We will analyze import demand functions. One phenomenon specific to import functions is that the domestic indexes that are typically used have weights that are quite unlikely to be adequate for analyzing competition between imports and domestic goods. For example, services substitute little if any with imports and should have low weight in the domestic price index. We hope from our analysis to identify the prices and incomes that are most critical in determining imports and also to obtain better estimates of the aggregate elasticities.