Module 3, Oct 14, Seminar and lab. The Ethics Case Brief is due at the start of class.
Data sovereignty and privacy
Ethics Case Brief deadline
Your Technology Ethics Case Brief (3-page brief plus 5-minute recorded walkthrough, 20% of your grade) is due at the start of class. Come knowing your case; the opening discussion is built from everyone’s briefs.
Background
Module 2 asked whether the classification is biased. This week asks who governs the data at all. At the start of class we sort your Case Brief cases by one criterion, whether the affected community could have refused, and almost every harm case (MiDAS, Allegheny, SafeRent) lands in the no-control column. The conceptual move runs from privacy, an individual defensive right, to data sovereignty, a collective governing right. You can consent to a survey, but your community cannot un-consent to what gets built from the aggregate.
Two of those cases carry numbers worth having at hand before the sort. MiDAS, Michigan’s automated unemployment-fraud system, cost $46 million to build, was wrong roughly 93 percent of the time in the fraud determinations it made without human review, falsely accused about 40,000 people, and figures in more than 11,000 bankruptcy filings; the class action, Bauserman v. Unemployment Insurance Agency, was allowed to proceed by the Michigan Supreme Court in 2022. SafeRent paid $2.28 million in 2024 to settle a Fair Housing Act suit over an AI tenant-screening score that penalized Black and Hispanic applicants using housing vouchers. Nobody scored by either system had a way to refuse the scoring.
This fight is underway now over immigrant data, and at a much larger scale. Georgetown’s American Dragnet investigation found that ICE has run face-recognition searches on the driver’s-license photos of roughly one in three US adults, can access license data on three in four, tracks vehicle movements in areas covering about 70 percent of the adult population, and can find a mover’s new home address through utility hookup records. The searches began early and quietly; ICE was running face recognition on Rhode Island DMV records by 2008, a decade before the practice became public. Georgetown puts ICE’s total spending on surveillance and data infrastructure at about $2.8 billion between 2008 and 2021.
The buying has since accelerated. In December 2025 ICE awarded skip-tracing contracts to 13 private firms under a vehicle worth up to $1.2 billion over two years, with contractors receiving as many as 50,000 names a month to locate using AI tools. That figure describes one 2025 contract vehicle; it sits on top of, and is distinct from, the $2.8 billion in cumulative spending above. The federal rule that would have treated data brokers as consumer reporting agencies was withdrawn on May 15, 2025, and EPIC documents ICE and CBP buying mobile-location, license-plate, and jail-booking data from brokers, including in places where local policy prohibits direct information sharing with ICE.
Key ideas
Privacy and data sovereignty
Privacy protects an individual from exposure; sovereignty gives a community authority over collection, use, and refusal. The difference is concrete in research ethics. A university IRB requires individual informed consent, while OCAP-style governance also requires collective consent, ongoing community control, and the right to withdraw the entire dataset. “We anonymized it” answers the privacy question and leaves the sovereignty question untouched.
Anonymization itself holds up worse than its reputation. A DW investigation of roughly ten billion location data points from the broker market found trails that revealed individual visits to psychiatric clinics, the kind of re-identification that “anonymized” location data rarely resists.
The immigrant-data pipeline in the Background shows the gap. An individual consents once, at the point of collection, and has no say in the resale and reuse that follow.
Individual privacy
Consent is given once, at the point of collection.
Collective sovereignty
Everything past that first handoff lies beyond any individual or community’s control. That gap between a one-time individual consent and an ongoing collective say is what data sovereignty names.
Networked privacy
Individual consent leaks even on its own terms, because data about one person describes others. McNealy works through two genetic cases: the Golden State Killer was identified through a public genealogy database that matched a relative’s DNA, and the 23andMe breach exposed people connected to account holders who had opted in. A person who never took a DNA test can still be found through a cousin who did, which is why McNealy treats community consent as working infrastructure for biodata governance and takes it seriously as neither a ceiling nor a floor.
The exposure falls unevenly by class. Madden, Gilman, Levy, and Marwick’s survey evidence shows low-income Americans are more dependent on mobile phones and use fewer privacy-protective tools, and their account of networked privacy harms, where a person is judged by their social network’s data rather than their own, runs through employment screening, higher-education access, and predictive policing.
Taylor’s data justice framework gives this cluster of problems a shared vocabulary: fairness in how people are made visible, represented, and treated as a result of the data they produce. Her three pillars are visibility, the right to be represented weighed against the right to informational privacy; disengagement, the freedom to stay out of data markets; and antidiscrimination. Disengagement is the pillar US law protects least, since there is no comprehensive federal data-privacy statute to exit through.
The four power questions
D’Ignazio and Klein’s first data-feminism principle asks of any data system: who does the work, who benefits, whose priorities become products, and who is harmed. The questions are built on Patricia Hill Collins’s matrix of domination, which tracks how power operates across four domains at once, structural, disciplinary, hegemonic, and interpersonal, so a system can be lawful in one domain and coercive in another. The privilege hazard sits behind these questions, the problems a data team from dominant groups cannot see because its members have never lived them.
The Allegheny Family Screening Tool gives each question a number. Carnegie Mellon researchers found it flagged 32.5 percent of Black children for mandatory neglect investigation against 20.8 percent of white children, caseworkers disagreed with the algorithm’s score about a third of the time, and a technical glitch fed workers wrong risk scores for over two years. Applied to collection specifically, the principle asks whether the people being counted had any say in the decision to count them; the families scored in Allegheny County had none.
CARE and OCAP in production
CARE (Collective Benefit, Authority to Control, Responsibility, Ethics) puts people and purpose alongside the data-centric FAIR principles; the drafters’ summary of the pairing is “Be FAIR and CARE.” The principles were drafted on November 8, 2018, at an International Data Week workshop in Gaborone, Botswana, by thirteen drafters from seven countries, building on the Māori data-sovereignty network Te Mana Raraunga, the US Indigenous Data Sovereignty Network, and the Aboriginal and Torres Strait Islander collective Maiam nayri Wingara, whose own five principles are control, contextuality, relevance, accountability, and protection.
OCAP (Ownership, Control, Access, Possession) is older and further into practice. It has been in force since 1998 and governs the First Nations Regional Health Survey, backed by community data-sharing agreements, and a 2025 five-phase framework now maps the road from no governance to fully Indigenous-led governance in practice. These regimes are useful precisely because they are enforceable, though scholars also warn that community consent does not substitute for individual consent.
US tribal nations supply the working examples closest to home. Carroll, Rodriguez-Lonebear, and Martinez document three: the Pueblo of Laguna partnered with the University of New Mexico to build its own census software while keeping ownership of the data, the Swinomish Tribe’s climate-change data agreements require tribal leadership approval before any external use or publication, and the Navajo Nation’s human research review board has required researchers to transfer their data to the Nation at project completion since 1996. A systematic review of 34 Australian studies finds the same machinery emerging there in four tiers of agreement, from international bodies down to individual contracts; accountability and self-determination are the most-cited principles in that literature, access and community voice the least.
The frameworks are now being extended to generative AI. Brookings argues that tribal nations need their own AI governance on the premise that data is kin, and lists current deployments: the Cherokee Nation’s legal-research agent working across tribal court decisions, the Morongo Band’s internal legal-reference chatbot, and the Yakama Nation’s AI-assisted irrigation. Each runs on community data, so each deployment carries a governance decision with it.
Refusal as a justice claim
Du Bois and Hull House used data to make injustice visible; communities under CARE sometimes refuse to be counted to stay safe. Counting and refusing are both justice claims about the same data, and what separates them is who holds authority over the collecting.
The cost of counting without that authority is documented in Bangladesh. After UNHCR’s 2018 biometric registration of Rohingya refugees, Human Rights Watch found that roughly 830,000 names with biometric data went to Myanmar, the government the refugees had fled, for repatriation-eligibility screening. Of 24 refugees HRW interviewed, one recalled being asked about sharing data with Myanmar; most had received English-only consent receipts with the yes boxes already checked, and 21 people whose names appeared on verified repatriation lists went into hiding out of fear of forced return.
UNHCR said publicly that the registration data was not linked to repatriation while using it for exactly that purpose, and it never completed a data protection impact assessment. A registration built to deliver aid became, through one data transfer, a screening list held by the government the refugees were fleeing. Te Hiku Media shows a third option beyond counting and refusing, being counted on your own license.
The Te Hiku Media case
Te Hiku Media, a Māori community media organization in Aotearoa New Zealand, ran the Kōrero Māori campaign to build speech recognition for te reo Māori and gathered about 300 hours of community-donated audio in ten days. The resulting model reached roughly 92 percent accuracy, giving the language a tool most minority languages still lack.
The governance move came next. Te Hiku released its work under a Kaitiakitanga (guardianship) license, restricting use to the benefit of Māori, and refused to sell when large tech companies asked. The community that donated its voices keeps authority over what those voices train. Control over a linguistic minority’s data shapes whether the language survives and whether its speakers can be understood in a clinic or a court.
The project has grown since the refusals. Papa Reo, the platform Te Hiku built on the donated corpus, states its mission as enabling smaller indigenous language communities to develop their own speech recognition and language technology, with work extending to Hawaiian and Samoan, and it is the only non-university recipient of a seven-year New Zealand Strategic Science Investment Fund grant, awarded in 2020. The license’s operating rule is that any benefit derived from the data flows back to its source. The IEEE Spectrum profile puts the recognizer at 92 percent accuracy for te reo Māori speech and 82 percent for bilingual speech.
The model of negotiated release is spreading. At the University of Waikato, Te Taka Keegan’s team built a te reo Māori text-to-speech voice from 7 hours and 45 minutes of recordings donated by a single speaker, Ngaringi Katipa, reaching a 6.78 percent word error rate. Rather than publish the voice openly, Keegan is negotiating release terms with the three iwi affiliated with the donor, Waikato, Maniapoto, and Raukawa, so the decision about a synthetic voice rests with the communities the voice belongs to. The same article tracks parallel projects elsewhere, including Michael Running Wolf’s FLAIR initiative for North American Indigenous-language speech models and the Barcelona Supercomputing Center’s Catalan text-to-speech system Matxa.
Te Hiku’s CEO, Peter-Lucas Jones, argues in a 2024 talk that guardianship licensing exists to keep large technology companies from extracting Indigenous language, sound, and story data for profit. The license is a working answer to the authority-to-control question the frameworks pose in the abstract.

Before class
Readings
- D’Ignazio & Klein (2020). Data Feminism, Chapter 1: “The Power Chapter”. MIT Press.
- Carroll et al. (2020). The CARE Principles for Indigenous Data Governance. Data Science Journal.
No pre-reading for the lab: the Digital Defense Playbook activities are run together in session.

In class
Seminar
We open with your Case Briefs, sorting each case by who controlled the data and whether the community could have refused. The seminar then takes up the Power Chapter’s four questions, CARE and OCAP as working governance, and the Te Hiku case, before the revote.
Lab: the Digital Defense Playbook
The lab runs popular-education activities from the Our Data Bodies project as written, starting with group agreements. You trace how one ordinary act (a transit swipe, a benefits form) becomes a data stream passing through many hands, run Blacklight on a site you use to see the trackers actually loaded on it, map your own “data body,” and then, using a fictional persona, prompt an LLM to assemble the kind of profile a data broker sells and interrogate the result with CARE, AI disclosure included. The lab closes with each person naming one concrete practice of collective data governance, from a community data-sharing agreement to a Kaitiakitanga-style license.
Further reading
- Global Indigenous Data Alliance. The CARE Principles for Indigenous Data Governance.
- Maiam nayri Wingara Indigenous Data Sovereignty Collective. The MnW principles.
- First Nations Information Governance Centre. The First Nations Principles of OCAP, and their video explainer.
- Carroll, S. R., Rodriguez-Lonebear, D., & Martinez, A. (2019). Indigenous data governance: Strategies from United States Native nations. Data Science Journal.
- Prehn, J., & Walter, M. (2023). Indigenous data sovereignty and social work in Australia. Australian Social Work.
- Prehn, J. (2025). Implementing Indigenous data sovereignty: A five-phase framework. Australian Journal of Social Issues.
- Trudgett, S., et al. (2022). A framework for operationalising Aboriginal and Torres Strait Islander data sovereignty in Australia. eClinicalMedicine.
- Brookings Institution. Avoiding the next digital divide: Defining digital sovereignty for Tribal Nations in the AI age.
- Winkless (2026). Māori Data Sovereignty Inspires New AI Voice Models. IEEE Spectrum.
- Jones, P.-L. (2024). Protecting our future: Indigenous data sovereignty (talk video). Australian Digital Alliance.
- Taylor (2017). What Is Data Justice? Big Data & Society.
- Madden, Gilman, Levy & Marwick (2017). Privacy, Poverty, and Big Data. Washington University Law Review.
- McNealy, J. (2024). Community consent: Neither a ceiling nor a floor. CHIRON.
- Wang, N., et al. (2022, updated 2025). American dragnet: Data-driven deportation in the 21st century. Georgetown Center on Privacy & Technology.
- EPIC (2025). How Data Brokers Harm Immigrants.
- EPIC (2025). EPIC condemns CFPB’s withdrawal of proposed rules to rein in data brokers.
- Human Rights Watch (2021). UN shared Rohingya data without informed consent.
- DW Documentary. (2026). Dangerous apps: In the web of data brokers (documentary).
- Crocker, T. (2025). Balancing data privacy and social good. International Social Work.