Gary Becker’s 1957 book-length monograph, The Economics of Discrimination, was one of the first economic treatments of discrimination in the marketplace, giving impetus for a new field of economic research. Among its many contributions, Becker’s work provided a theoretical framework to quantify non-pecuniary motives in discriminatory behavior within labor markets. Until then, such motives were the provenance of sociology, psychology, and anthropology. What could economics possibly have to say about issues not relating to money?

Plenty, as Becker revealed not only in his book but throughout his influential career, as he explored the many applications of economic analysis to human behavior. On discrimination in hiring, Becker identified three potential biases: managerial, co-worker, and customer. Economists have since investigated these biases, both theoretically and empirically, to determine their influence on hiring. However, five decades of measuring and identifying discriminatory hiring patterns have not resolved all quantitative challenges. For example, limited data mean that researchers must rely on aggregated measures that hide effects related to the biases of managers, employees, and customers. As Becker described, to thoroughly parse the role of discrimination in hiring, we need to consider those distinct discriminatory preferences.
This paper addresses that gap by introducing a theoretical framework that shows that to disentangle managerial and customer biases, one’s data must possess certain features:
- an objective metric for assessing worker quality,
- detailed information regarding the composition of hiring decision-makers, and
- a measure of customer bias magnitude.
It turns out that an ideal setting for this investigation is Major League Baseball’s (MLB) annual draft, which is how MLB assigns amateur players (whether high school, college, or amateur baseball clubs) to its professional teams. This setting adheres to an external validity litmus test for data developed by one of this paper’s co-authors, UChicago’s John A. List (see “In defense of external validity”). Based on List’s rationale from his 2020 paper, the authors conclude that it is nearly impossible to find a more appropriate setting that allows for rigorous testing of the above framework.
In defense of external validity
To test their theoretical insights (models) against reality (empirics), economists and other social scientists employ data. These data are often limiting in their explanatory value as they necessarily represent a fraction of a possible dataset. In the present case, for example, to include every piece of data regarding every hiring decision in an economy over time would be impossible, not to mention unwieldy. The trick, rather, is to gather data from a particular subset that delivers strong findings from which we can extrapolate to the rest of the world.
This is more than an academic exercise, as the point of most social science is to impact policymaking. And if we are going to affect policymaking—that is, directly impact people’s lives—we want to ensure that our data are valid beyond our subset; that is, our data should be externally valid. For some, such validity is a humbug. For these skeptics of empirical economics, even the slightest doubt about validity renders a study moot.
In “Non Est Disputandum De Generalizability? A Glimpse into the External Validity Trial,” a satirical (and, rare to say for an economics paper, entertaining) defense of external validity, UChicago’s John A. List argues that it is possible to pass an external validity test. Indeed, unique empirical settings—in our case here, MLB hiring practices—are not always a distraction from reality; rather, when that uniqueness allows for relevant testing that no other setting can achieve, then a level of “perfection” is possible whereby we can confidently generalize (and scale) to the rest of the world.
Speaking of scaling, this little article only begins to describe List’s longer argument for the validity of empirical research, and the reader is encouraged to visit the full paper via the link above. That said, in sum, here are List’s four tenets of empiricism necessary to address external validity:
1. Theory and empiricism are symbiotic: theory provides a structure for thinking about the world, empirical work tests whether that structure is approximately correct and informs future theories.
2. One swallow does not make a summer: each study moves priors by an amount corresponding to its quality and the strength of priors.
3. To explain differences in observed choices across settings, ask if preferences, constraints, or beliefs have changed.
4. Uniqueness of a setting can be a key strength, not a weakness, if it isolates a particular channel or causal mechanism effectively.
The authors examine drafting (or hiring) decisions made by MLB teams from 2008 to 2019, when about 12,000 players were drafted, including scouting evaluations and detailed information about each scout, including racial background. (Baseball scouts evaluate players for MLB teams, including on location during games and at training facilities; think of them as a traveling HR department.) Publicly available data on thousands of players—drafted and undrafted—allow the authors to construct the first key metric necessary to distinguish sources of discrimination: an objective measure of player quality. In other words, all players deemed high quality should be drafted, regardless of race or other discriminatory factors.
To fulfill the second dataset described above—the racial composition of hiring managers—the authors also collect comprehensive data on all MLB scouting directors, including their racial backgrounds. These data allow the authors to assess whether scouting directors exhibit a propensity to recruit players of their own race.
Finally, to measure customer (fan) bias, the authors study a naturally occurring event, the Black Lives Matter (BLM) movement in June and July 2020, during which all MLB teams posted messages on social media. The authors then perform a textual analysis of responses to such postings to create an index of fan bias. Further, the authors analyze stadium attendance data from 2008-2019, examining its correlation with the racial composition of the team.

Thus, armed with data addressing worker quality, manager discrimination, and customer bias, the authors apply these empirical insights to their models to find the following:
- There is no significant association between race and the likelihood of a player being drafted. When controlling for prospect quality, African American players exhibit a slightly higher draft probability compared to their White counterparts.
- Player compensation is generally consistent across racial groups. However, controlling for player quality, there is some evidence suggesting that Asian and White players receive lower signing bonuses compared to their peers.
- That said, patterns of discrimination loom deep within the data. There is a strong correlation between the drafting of African American players and customer bias during the early rounds of the draft: fan bias is associated with whether a player of a certain race is drafted early. This fan bias correlation, however, is reduced in the later rounds. Collectively, these findings suggest that MLB clubs are likely considering customer preferences when selecting players who will attract significant scrutiny and public attention.
- Conditional on player quality, scouting directors demonstrate a bias toward players of their own race, with these players 38 percent more likely to be drafted later (during rounds 26-40). This suggests that when the stakes are lower and public scrutiny is reduced, scouting directors are more inclined to express their personal preferences, which is supported by the low probability that these players will reach the major leagues.1
- Finally, and related to the above finding, these revealed biased preferences carry economic costs: Teams draft lower-valued players when fan bias increases. While such customer bias bears significant opportunity costs (measured as reduced number of wins per season), the financial impact of managerial bias is limited, though, as these players are long shots to reach the majors.
Bottom line: The authors’ novel theoretical and empirical combination provides a framework for analysis of discrimination in economic settings where multiple sources of bias interact simultaneously, including biases hiding within aggregate measures. Likewise, and importantly, the authors’ results plausibly generalize to other markets; that is, this work adheres to List’s external validity test.
For scholars, this means caution when examining data for discrimination using establishment level data, as they run a risk when mining findings from aggregate data. When there is tension in biased preferences between management and customers, key aggregates can underestimate, or mask, key biases.
For policymakers, understanding the exact channels of bias is key to developing effective and scalable interventions, and this work offers a framework for modeling and estimating relevant sources of discrimination.








