Skip to main content icon/video/no-internet

Maximum likelihood estimation (MLE) is one of the most commonly used estimation methods. Some features that prediction methods should have are efficiency, unbiasedness, and minimum variance. The method of maximum likelihood (ML) has all of these features and is a very reliable estimator. MLE is also the basis of many statistical methods. In some cases, the likelihood function is mathematically difficult to compute. Nevertheless, parameters that are suitable for many distributions can be estimated with the MLE. This entry discusses general aspects of MLE as well as the ML estimators and their properties.

General Aspects of MLE

MLE was developed by R. A. Fisher in the 1920s. Its main aim is to make observed data most likely, that is, to maximize the likelihood function. MLE has an important place in the theory of many statistical methods. Furthermore, many inference methods in statistics are based on MLE, including chi-square test, Bayesian methods, inference with missing data, modeling of random effects, Akaike information criterion, and the Bayesian information criteria.

One of the most important features of the ML method is that the estimators obtained are asymptotically unbiased and efficient. An important disadvantage of the method is that mathematical solutioning of some maximization problems can be difficult in the estimation process.

One of the most reliable ways to predict variables is the MLE. The parameter values most likely to produce the observed data are selected from the MLE. Based on numerous scientific studies, the MLE provides an optimal and reliable estimate for many problems encountered.

The Fisher information matrix appears in MLE as a measure of independency between estimated parameters. The inverse of the Fisher information matrix gives the covariance matrix for the estimation errors of the parameters. As a result of this feature, the orthogonalization of the parameters ensures that the estimates of the parameters are distributed independently from each other. However, diagonalizing the Fisher matrix by linear algebra locally or at specific parameter values does not make sense because the parameter values are unknown, that is, they have not yet been estimated.

The MLE is a reasonable choice for an estimator because it is the parameter point for which the observed sample is most likely. There are two inherent disadvantages associated with the general problem of finding the maximum of a function, and therefore, MLE. The first problem is that of actually finding the global maximum and verifying that, indeed, a global maximum has been found. In many cases, this problem reduces to a simple differential calculus exercise, but, sometimes even for common densities, difficulties do arise. The second problem is that of numerical sensitivity—that is, how sensitive is the estimate to small changes in the data? This problem is a mathematical rather than a statistical problem associated with any maximization procedure. Unfortunately, sometimes a slightly different sample will produce a vastly different MLE.

The following is a real-life example to understand the MLE process. Suppose that four text messages arrive on a person’s phone in the morning. However, the person deleted one of the messages carelessly before reading it. If two of the remaining three messages are about monthly bills, what might be a good estimate of k, the total number of monthly bills among the four messages? Frankly, k must be two or three.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading