Motivation: Applications of Regression in Research#
Regression analysis is a powerful statistical tool used in various research fields to understand the relationship between variables and make predictions. Regression analysis is also the fundamental tool of Econometrics, the discipline in economics that studies specific statistical tools to examine economic phenomena.
Regression analysis enables a series of interesting applications in quantitative research. In this introductory section, we will review some of these, including data extrapolation, causal relationship inference, and the construction of predictive models.
Uses of Regression#
1. Data Extrapolation and Prediction:
A first application of regression is data extrapolation. Let me begin by illustrating this use with a case study that I find particularly compelling: how Xbox gaming data was used to predict the outcome of the 2012 U.S. presidential election [Wang et al., 2015].
As you may know, traditional polling methods, which typically rely on telephone surveys, have become increasingly less effective due to declining response rates and biased samples. Polls can easily be skewed toward specific demographics. Recognizing this challenge, researchers proposed a novel approach: leveraging the vast Xbox user base to launch a survey and create a dataset that reflected the demographic characteristics of the actual voting population. At first glance, this might sound strange: Xbox users are expected to be biased relative to the general population, specifically toward younger people!
However, in this case, researchers employed an ingenious strategy to address the bias. Taking advantage of the thousands of responses they could obtain from Xbox users, they constructed a grid of different demographic profiles (e.g., age groups, gender, etc.) to estimate the voting tendencies of each group based on their characteristics. When data for specific combinations of characteristics were unavailable or too sparse, the researchers used interpolation, drawing on the results of similar demographic groups. Regression played a crucial role in this interpolation process, allowing researchers to make predictions even for groups with limited data. Once the grid of voting tendencies for each group was complete, the researchers weighted each group according to its share in the voting population (a process commonly referred to as post-stratification). As a result, they were able to obtain predictions quite close to the actual election outcomes.
2. Inference of Causal Relationships:
Beyond prediction, regression plays a crucial role in inferring causal relationships between variables. Take, for example, how technology companies like Netflix evaluate the causal impact of their product decisions.
Suppose Netflix’s product team wants to know whether a new personalized recommendations feature increases subscriber retention. A first idea would be to compare the retention of users who use that feature with that of users who do not. But, as we will see when we study causal inference, this type of comparison can be misleading: users who actively choose to use a new feature are likely more engaged with the service in the first place. Comparing “apples to oranges”—users who adopt the feature versus those who do not—would conflate the effect of the feature with pre-existing differences between the groups.
To address this problem, Netflix randomly assigns some users to the group that receives the new feature (the treatment group) and others to the group that does not (the control group). This random assignment ensures that both groups are comparable before the experiment. Using regression models, researchers estimate the causal effect of the feature on retention, controlling for the small differences that may exist between groups even in a well-designed experiment. This approach—so-called randomized controlled experiments, or A/B tests—is today the standard in the technology industry for making product decisions grounded in causal evidence.
We will see that regression is not only useful for analyzing randomized experiments. It is also the central tool for drawing causal inference from observational data—where random assignment is impossible—through techniques such as instrumental variables, differences-in-differences, and regression discontinuity designs.
3. Construction of Predictive Models:
As a third use case for regression methods, regression is widely employed to build predictive models, allowing us to estimate the value of a “dependent” variable based on the values of “independent” variables. Consider, for example, how large technology companies determine their employees’ salaries.
Companies like Google, Meta, or Amazon manage thousands of employees in very different roles. To set salaries consistently and equitably, they build regression models that predict an employee’s expected salary as a function of their observable characteristics: years of experience, technical specialization, seniority level, geographic location, and performance-evaluation results. This type of model—known internally as a “compensation model”—works much like the hedonic pricing model in economics: it decomposes salary into the individual contributions of each employee characteristic. By applying regression, companies can estimate how much each additional year of experience is “worth,” how much of a premium is paid for working in San Francisco versus Buenos Aires, or how much an exceptional performance review affects pay. This information is then used to project the salary of new hires and to verify that internal compensation is consistent with the market.
While regression models typically assume linear relationships between variables, regression can also be used to model nonlinear relationships. By applying transformations to the variables—such as the logarithm of salary or of years of experience—researchers can capture more complex relationships between characteristics and compensation, improving the model’s accuracy.
These examples demonstrate the versatility of regression analysis. From predicting election outcomes to inferring causal relationships in product experiments and building predictive models for compensation management, regression provides a robust framework for understanding complex phenomena and extracting valuable insights from data.
References#
Wei Wang, David Rothschild, Sharad Goel, and Andrew Gelman. Forecasting elections with non-representative polls. International Journal of Forecasting, 31(3):980–991, 2015.