UMCubed All articles
Math Education

Coded Verdicts: The Hidden Mathematics Shaping Who Goes to Prison in America

UMCubed
Coded Verdicts: The Hidden Mathematics Shaping Who Goes to Prison in America

Photo: U.S. Air Force photo by Senior Airman Ashley Talley, Public domain, via Wikimedia Commons

When Numbers Become Verdicts

In courtrooms from Wisconsin to California, a defendant's fate may hinge not only on testimony and evidence, but on a score—a single number produced by a proprietary algorithm. Tools such as COMPAS (Correctional Offender Management Profiling for Alternative Sanctions) and the Public Safety Assessment assign numerical risk ratings that judges consult when determining bail, sentencing length, and parole eligibility. To many observers, a mathematical score implies objectivity. To researchers who have examined these systems closely, the reality is far more troubling.

The promise of algorithmic decision-making in criminal justice was straightforward: replace inconsistent human judgment with consistent, data-driven analysis. What emerged instead was a system that encodes historical inequities into mathematical form—and then presents those inequities as neutral fact.

The Architecture of Bias

To understand how bias enters these models, one must first understand what they are measuring. Most criminal risk-assessment tools are built on logistic regression or decision-tree frameworks trained on historical criminal justice data. The models assign weights to input variables—prior arrests, age, employment status, residential zip code, family history of incarceration—and produce a probability estimate of future criminal behavior.

Here lies the fundamental problem. When a model is trained on data generated by a justice system that has historically over-policed Black and low-income communities, the historical record itself is distorted. Prior arrests, for instance, reflect not only actual behavior but also the intensity of law enforcement presence in a given neighborhood. A young Black man in a heavily surveilled urban zip code accumulates a different arrest history than a white peer who engages in identical behavior in a less-monitored suburb. When that disparity enters a training dataset, the model learns to treat it as a meaningful signal. The mathematics then amplifies what was always a sociological artifact.

A landmark 2016 investigation by ProPublica analyzed COMPAS scores assigned to more than 7,000 people in Broward County, Florida. Researchers found that Black defendants were nearly twice as likely as white defendants to be falsely flagged as high risk for future crimes. White defendants were more frequently misclassified as low risk despite subsequently reoffending. The algorithm's overall accuracy was roughly 65 percent—a figure that would be unacceptable in virtually any other high-stakes scientific context.

Fairness Is a Mathematical Concept—and a Contested One

What makes algorithmic bias in criminal justice particularly complex is that different definitions of mathematical fairness are genuinely incompatible. Researchers distinguish between several fairness criteria that cannot all be satisfied simultaneously.

Calibration requires that a risk score of, say, 70 percent corresponds to a 70 percent actual recidivism rate across all demographic groups. Equal false positive rates require that individuals who will not reoffend are incorrectly flagged at the same rate regardless of race. Equal false negative rates require that individuals who will reoffend are missed at the same rate regardless of race.

A 2016 paper by Chouldechova demonstrated mathematically that when base rates of recidivism differ between groups—as they do in the United States, largely due to socioeconomic factors—it is arithmetically impossible to satisfy calibration and equal error rates simultaneously. This is not a software flaw. It is a theorem. Any algorithm operating on unequal underlying distributions must make a choice about which group bears the cost of its errors. Currently, that cost falls disproportionately on defendants of color.

Students learning probability and statistics should encounter this tension directly. The fairness impossibility result is one of the most consequential mathematical findings of the past decade, yet it rarely appears in undergraduate curricula.

Auditing the Black Box: What Mathematical Literacy Demands

For those who wish to scrutinize these systems, the required knowledge spans several mathematical domains. Logistic regression and Bayesian inference form the statistical foundation. Linear algebra underlies the weight matrices that transform input features into scores. Concepts from information theory—particularly conditional entropy—help analysts measure how much a protected attribute like race is implicitly captured by ostensibly neutral proxy variables.

Perhaps most critically, auditors must understand the difference between correlation and causation in model training. A variable that predicts recidivism in a dataset is not necessarily causally related to recidivism; it may simply correlate with the conditions that make re-arrest more likely. Untangling these relationships requires causal inference methods that are only beginning to enter standard statistics curricula.

Several states have begun requiring algorithmic impact assessments before deploying risk tools in criminal proceedings. Organizations such as the Markup and the Stanford Computational Policy Lab have developed open-source audit frameworks. However, these efforts are constrained by a persistent shortage of mathematically trained professionals who also possess the legal and ethical vocabulary to interpret their findings within a justice context.

A Mandate for the Next Generation

American mathematics education has long prioritized technical proficiency over applied civic reasoning. Students learn to solve differential equations but rarely to question the social systems those equations model. The emergence of algorithmic governance—in criminal justice, credit scoring, hiring, and healthcare—demands a different orientation.

Universities and community colleges have an opportunity to design courses that bridge quantitative methods and policy analysis. Programs in applied mathematics, data science, and statistics should incorporate case studies from criminal justice alongside the technical content. Law schools, conversely, should require quantitative literacy sufficient for attorneys and judges to interrogate algorithmic evidence.

Some institutions are moving in this direction. The University of California, Berkeley's Human Contexts and Ethics program embeds social analysis into data science training. Similar initiatives at Carnegie Mellon and MIT treat algorithmic accountability as a technical discipline, not merely a philosophical add-on.

Mathematicians who engage with these questions are not abandoning rigor for advocacy. They are extending the tradition of applied mathematics into one of the most consequential domains of contemporary life. The geometry of justice is not a metaphor. It is a set of real mathematical choices with real consequences for real people—and understanding those choices is precisely the kind of work that defines a scientifically literate society.

For students considering careers in data science, statistics, or policy analysis, the message is direct: the most important problems of the coming decades will not be solved by technical skill alone. They will require the courage to ask what the numbers are actually measuring, and the literacy to answer that question with precision.

All Articles

Related Articles

When the Shelves Go Empty: The Combinatorial Mathematics Keeping America's Supply Chains Alive

When the Shelves Go Empty: The Combinatorial Mathematics Keeping America's Supply Chains Alive

Signals from the Mind: The Linear Algebra Powering the Next Frontier in Brain-Computer Interface Research

Signals from the Mind: The Linear Algebra Powering the Next Frontier in Brain-Computer Interface Research

Clicks, Clusters, and Cascades: The Mathematical Architecture Behind What Goes Viral

Clicks, Clusters, and Cascades: The Mathematical Architecture Behind What Goes Viral