Module 2, Sep 16, Seminar and lab. The Ethics Case launches this week.

Algorithmic bias and fairness

Background

The harm side of algorithmic bias gets its strongest evidence in the MiDAS case below. The override question has its own direct evidence. At Allegheny County’s Family Screening Tool, a CHI study found that caseworkers who overrode the algorithm’s risk score cut the Black-white screen-in disparity from 20 to 9 percent, drawing on context the algorithm lacked, such as a family’s engagement with private services that never enter administrative records. A follow-up analysis on updated county data confirmed the finding.

The case against a standing override is institutional. Bovens and Zouridis describe systems like this one as built specifically to move judgment from the caseworker to the system’s designer, and returning that judgment to case-by-case street-level discretion, Lipsky’s term for it, risks reintroducing the inconsistency the system was meant to remove. The Allegheny tool’s independent evaluation states that tension directly: a risk score can systematize otherwise-inconsistent human screening decisions, and by construction it draws on records that disproportionately exist for poor and minority families.

Automation bias adds a practical worry on top. A systematic review of 35 studies finds people defer to automated recommendations even when they know better, and explanation features do not reliably fix it, so an override power that exists on paper is not the same as one that gets used when it matters.

The public has been polled on questions close to this one. In a 2018 Pew Research Center survey of 4,594 US adults, 56 percent said using algorithms to make criminal risk assessments of people up for parole was unacceptable, and the most common objections were that a program cannot capture how individuals and circumstances differ and that people can change; in the same survey, 58 percent said computer programs will always reflect some level of human bias. In a Pew survey of 11,004 adults fielded in December 2022, 71 percent opposed letting AI make a final hiring decision and 66 percent said they would not want to apply for a job where AI helped make the hiring decisions.

Key ideas

Four sources of algorithmic bias

Algorithmic bias has four sources. Training data encodes historical inequity: a model trained on past decisions learns the record, including the parts of the record that exist only because some families get reported more. Anonymous and mandated reporters flag Black and biracial families roughly 3.5 times more often than white families, so any child-welfare model trained on referral data starts from that ratio.

Proxy variables stand in for what a system cannot measure directly. In Benjamin’s Science example, a widely used health algorithm took cost of care as a stand-in for health need, and because providers had historically spent less on Black patients, equally sick Black patients scored as healthier. In Allegheny County, re-referral to the hotline stands in for maltreatment itself. No biased intent is required in either case; a badly chosen proxy does it alone.

Feedback loops close the circle: surveilled communities generate more data, which raises their scores, which brings more surveillance. New York City’s child-welfare scoring tool rates families on 279 variables, including neighborhood and the mother’s age, in a system where Black families are already referred at seven times the rate of white families and Black children are 13 times more likely to be removed; families and their lawyers are never told a case was algorithmically flagged, so no one outside can inspect the loop. The fourth source, automation bias, is the human habit of deferring to a score even against one’s own better judgment, and it carries the other three into casework decisions.

The New Jim Code

The New Jim Code is Ruha Benjamin’s name for technology that reproduces racial hierarchy behind a veneer of neutrality and efficiency, where “objective” and “data-driven” function as an alibi. The health-cost algorithm above is her signature case because nobody built it to ration care by race; it was designed to find patients who needed extra help, and the discrimination arrived inside the proxy.

In Noble’s Algorithms of Oppression, the case is search: Google results for “Black girls” returned demeaning content while the ranking presented itself as neutral relevance, so discoverability itself carried the bias. Kalluri’s corrective is to stop asking whether an AI system is fair and ask how it shifts power: who benefits, who is exposed to risk, and whether affected communities shape the systems used on them.

Technological benevolence is the helpful tool that still surveils.

Photograph of Ruha Benjamin speaking at a podium
Ruha Benjamin presenting at Data & Society’s Databites speaker series, October 2019. Her book Race After Technology coined the term “the New Jim Code,” this week’s required reading. Photo: Data & Society Research Institute, via Wikimedia Commons, CC BY 3.0.

The Allegheny Family Screening Tool

Since 2016, Allegheny County has scored child-neglect calls with a tool that draws on SSI records, mental-health diagnoses, jail records, Medicaid records, and zip code to produce a 1-to-20 risk score for the family named in the call. The design tension the Background describes comes from the county’s own commissioned evaluation, a study the tool’s critics and defenders both cite.

The AP investigation that later drew Department of Justice attention reported a Carnegie Mellon analysis showing the tool flagged 32.5 percent of Black children for mandatory investigation versus 20.8 percent of white children, that social workers disagreed with its risk scores about a third of the time, and that a technical glitch fed workers incorrect scores for more than two years without being caught. DOJ civil-rights attorneys began scrutinizing the tool in 2023 after complaints that it could harden bias against families with disabilities, and Oregon dropped a similar tool over racial-equity concerns.

An ACLU and Human Rights Data Analysis Group audit read the tool’s design choices as de facto policy: the score draws on juvenile-probation and behavioral-health records that overrepresent poor and Black families, and displaying only a household’s single highest child score could flag roughly 33 percent of Black households as high risk versus about 20 percent of non-Black households. HRDAG’s statistical work for the ACLU also confirmed the score penalizes parents with disabilities, since a mental-health diagnosis, even ADHD, or SSI eligibility feeds a score a family has no way to clear.

The override evidence sits on top of all this. Caseworkers who could see and overrule the score cut the Black-white screen-in disparity from 20 to 9 percent, and an ethnographic study of algorithmic tools in child welfare found caseworkers doing constant repair work around system malfunctions under time pressure, a process harm that degrades decisions independent of the algorithm’s statistics.

Competing definitions of fairness

ProPublica and the vendor Northpointe both ran the numbers on the COMPAS recidivism score, reached opposite verdicts, and both did their math correctly. ProPublica’s analysis of more than 7,000 people arrested in Broward County found Black defendants falsely labeled future criminals at 44.9 percent versus 23.5 percent for white defendants, while white defendants who went on to reoffend had been labeled low risk at 47.7 versus 28.0 percent; overall accuracy was 61 percent. The vendor’s defense was calibration, that a given score predicts about the same reoffense rate whichever group it lands on.

Kleinberg, Mullainathan, and Raghavan proved that these fairness definitions cannot all hold at once when base rates differ across groups, outside narrow special cases. Somebody has to pick, and picking a fairness metric means picking whose error the system tolerates, an ethics question social workers should insist on making explicit.

The choice surfaces wherever scores allocate scarce things. In 2024, SafeRent agreed to pay $2.2 million to settle claims that its tenant-screening score penalized housing-voucher holders by ignoring the voucher when assessing ability to pay, while over-weighting credit histories that reflect historical inequities; the settlement requires it to drop the score for voucher holders and have future scores validated by a third party. The fairness explorer lets you move the cutoff on two groups and watch the three tests trade off against each other.

Audits that changed systems

The record also holds audits that produced measured fixes. In Gender Shades, Buolamwini and Gebru tested three commercial gender classifiers, from IBM, Microsoft, and Face++, on a benchmark balanced by gender and skin type, and found error rates of up to 34.7 percent for darker-skinned women against a maximum of 0.8 percent for lighter-skinned men. The companies received the results privately in December 2017, and the study went public that February.

Raji and Buolamwini’s follow-up audit reran the benchmark seven months after the disclosure. All three companies had shipped new versions, IBM’s within 66 days; error on darker-skinned women fell by 17.7 to 30.4 percentage points, Microsoft’s gap between its worst and best subgroups shrank from 20.8 points to 1.5, and IBM’s from 34.4 to 16.7. Overall error fell at every audited company too, so narrowing the gap cost nothing in average accuracy. On the same benchmark in the same follow-up, Amazon and Kairos, which the original audit had not named, showed error rates on darker-skinned women of 31.4 and 22.5 percent.

Regulation has tried to make auditing routine, with thinner results so far. New York City’s Local Law 144, in force since July 2023, is the first US law requiring annual independent bias audits of automated hiring tools, with reports publicly posted; when 155 trained investigators checked 391 New York City employers, 18 had posted an audit report and 13 the required applicant notice. Because each employer decides for itself whether the law covers its tools, a missing report proves neither compliance nor violation; the audit that moved vendors had named its targets, published a benchmark anyone could rerun, and come back on a known date.

The MiDAS case

Michigan spent $47 million on MiDAS, an automated system meant to catch unemployment-insurance fraud. From October 2013 to September 2015 it made fraud determinations with no human review at all, quintupling fraud accusations relative to the prior manual process, while a 400 percent penalty structure pushed the agency’s fraud-penalty collections from about $3 million to over $69 million. It garnished wages and seized tax refunds without a judge ever being involved.

The determinations were wrong about 93 percent of the time, and roughly 40,000 people were falsely accused; at least 11,000 filed for bankruptcy. An internal state review that examined nearly 21,000 computer-flagged cases from a 22-month window confirmed the 93 percent figure, and the same reporting followed one claimant who had more than $14,000 garnished and was told he still owed $70,000, most of it penalties.

Accountability ran through Bauserman v. Unemployment Insurance Agency, which reached the Michigan Supreme Court in 2022 before falsely accused workers won the right to sue; a $20 million settlement covering roughly 3,000 class members was approved in January 2024, more than a decade after the system launched. The harm entered at more than one point: the model, the data, the removal of human review, the due-process-free collection design, and the accountability vacuum all qualify.

Michigan was not an outlier. In 2024 a federal judge ruled that TennCare Connect, a $400 million eligibility system in Tennessee, had illegally cut Medicaid coverage for thousands of people, and its builder made similar systems in more than 20 states. In 2020, a Dutch court struck down SyRI, the Netherlands’ welfare-fraud detection system, on human-rights grounds, finding it insufficiently transparent and verifiable and fed with data drawn disproportionately from poor neighborhoods.

Ethics Case Brief launch

The Ethics Case Brief (20% of your grade) opens this week and is due at the start of Week 7: pick a case, harm or success, and run the protocol we demo on MiDAS, in a structured brief of 3 pages max plus a 5-minute recorded walkthrough. Every AI-surfaced fact must be verified against a primary source. This week’s lab is the supervised rehearsal.

The Los Angeles prevention case

Los Angeles County’s Homelessness Prevention Unit uses a risk model on administrative data to find people likely to lose housing and reach them with cash aid and case management before the crisis. In the pilot phase, 335 enrollees were compared with 1,285 similar high-risk people who were not enrolled, and enrollees were 71 percent less likely to enter shelter or street outreach within 18 months. The researchers state that the pilot comparison is not yet causal, and a randomized trial is underway.

The program’s numbers are small next to the systems above. It had served 1,498 people, with households receiving an average of $6,469 in flexible assistance, and 86 percent retained housing on completing the program. Participation is opt-in and the model was equity-audited. Nothing about the math forced MiDAS’s direction; deployment decisions did.

Screenshot of the California Policy Lab report page on the Los Angeles County Homelessness Prevention Unit
The California Policy Lab’s publication page for its evaluation of the Homelessness Prevention Unit, the lab packet’s hopeful case. Screenshot of California Policy Lab, July 2026.

Before class

Readings

Screenshot of the Polity Books catalog page for Race After Technology by Ruha Benjamin
The publisher’s page for Race After Technology: Abolitionist Tools for the New Jim Code, this week’s required reading. Screenshot of Polity Books, July 2026.

In class

Seminar

The seminar covers the four sources of bias and the New Jim Code, then runs the sociotechnical protocol live on MiDAS as the worked example of the Case Brief method. We close with the COMPAS fairness dispute in plain language and launch Assignment 1.

Lab: the paired-case lab

You get two printed packets. Both describe a risk-scoring model built on administrative data, one used to flag families for investigation and the other to reach people with preventive aid. Packet A is Allegheny County’s Family Screening Tool, which scores families for neglect investigation, oversamples families experiencing poverty, and drew Department of Justice scrutiny; a CHI study found that caseworkers overriding its scores cut the Black-white screen-in disparity from 20 to 9 percent. That override is street-level discretion, Lipsky’s term for the judgment frontline workers exercise case by case, reasserting itself against exactly the kind of system Bovens and Zouridis describe as designed to move that judgment from the caseworker to the system’s designer. Packet B is LA County’s Homelessness Prevention Unit, above. Groups run the protocol on both cases, then each group names the single design or deployment decision that most separates harm from hope; almost none land on “the tech.” The class then revotes.

Same model class

scores risk from administrative data about vulnerable people

Surveillance register

Allegheny Family Screening Toola high score triggers an investigation

Care register

LA County Homelessness Prevention Unita high score triggers an offer of help

The statistical method is the same. The deployment decisions are what separate harm from help.

Further reading