Module 2, Oct 7, Seminar and lab. The partner introduction memo is due this week.

AI adoption in organizations

Background

Randomized trials keep finding that generative AI helps the least-experienced workers most, on exactly the kind of writing work intake involves, and a 2026 RCT found a similar pattern along education lines: with AI access, an education-based productivity gap of 0.548 standard deviations fell to 0.139, closing about three-quarters of it, and participants kept part of the gain in a follow-up round without AI. Automated systems in this same administrative space, including one run by this state, have hurt people at scale.

One reading of the moment says both sides overestimate the clock. Narayanan and Kapoor argue AI is a normal technology whose effects arrive over decades of diffusion: generative AI had reached 40 percent of US adults by August 2024 yet accounted for only 0.5 to 3.5 percent of work hours, a gap they compare to the roughly 40 years between electrification and measurable productivity gains. They also argue that safety bottlenecks, more than capability limits, govern how fast real deployment happens. On that view, the radical element of the motion is its 12-month deadline.

The sector is not waiting for governance to catch up. A recent benchmark survey found 92 percent of nonprofits already using AI while only 4 percent have documented workflows and 81 percent of use is individual and ad hoc. A Stanford HAI and Project Evident survey found 66 percent of nonprofits using some form of AI while 78 percent lack any formal organizational AI policy, even as about three-quarters of both nonprofits and funders believe their organizations would benefit from more AI in mission work. A national survey of social workers found 63.5 percent already using AI tools, mostly for correspondence, reports, and documentation, with almost no professional guidance. Across 11 federal agencies, GAO counted generative-AI use cases growing from 32 to 282 between 2023 and 2024, with the VA automating parts of medical imaging and HHS extracting outbreak information from publications, and with data-privacy compliance among the most-cited challenges. The adoption decision is being made by drift, one ad hoc prompt at a time.

Key ideas

Adopt, defer, or refuse

Deployment is a decision about the agency as much as the tool. The heuristic we build in seminar runs on a short list of questions:

  • If the output is wrong, who is harmed, and can the harm be undone?
  • Can a caseworker verify the output before anything acts on it?
  • Does the task determine someone’s rights or benefits?

Run a real decision through the adopt, defer, or refuse tool to see which of your answers push toward waiting or declining.

  • Whose data is involved, and did the client consent to this use?
  • Can the agency hold the vendor accountable, and can it leave?

The ground between wholesale adoption and refusal has published guidance. Perron and colleagues argue against treating AI adoption in human services as simply good or bad, proposing instead a risk assessment that varies by use case and context, across dimensions that include displaced professional judgment, bias, environmental impact, and labor exploitation. They flag local, on-premises language models as a lower-risk entry point and recommend starting with lower-risk applications before scaling, which is what adopting with guardrails looks like in practice.

Public agencies have sector guidance too. APHSA, the membership association of state and local human-services agencies, has published a posture statement on AI in human services whose first tenet treats AI as support for the workforce and keeps critical thinking and oversight with people. The association backs it with a working series on AI-assisted SNAP modernization and formal comments on federal AI rulemaking.

Most of these questions are really about discretion. Lipsky named the frontline workers who ration public services street-level bureaucrats. Their case-by-case judgment about how a rule applies to the person in front of them, exercised under caseload and resource pressure, is where policy actually gets made. Bovens and Zouridis describe what automating that judgment does. In what they call a system-level bureaucracy, discretion moves from the caseworker to whoever designed the system, before a client ever reaches the queue, and the people writing the decision rules are programmers and analysts no client will ever meet.

Some answers point to adopting with guardrails, others to deferring until the governance exists. Rights-determining eligibility decisions, irreversible actions, unverifiable outputs, and unconsented uses of client data sit in the refuse column, and MiDAS and TennCare show what deploying into that column costs.

Arkansas is the case closest to casework. When the state replaced nurse-led assessments with an algorithm that set Medicaid home-care hours, Tammy Dobbs, who has cerebral palsy, had her weekly hours cut from 56 to 32; a coding error had failed to account for diabetes complications in cerebral palsy patients, and one amputee was recorded as having no foot problems. A Center for Democracy and Technology review of a decade of litigation over algorithm-driven benefits cuts documents tools that cut care hours nearly in half and the legal theories that eventually won, from due-process notice requirements to the ADA. Undoing this kind of harm takes years of lawsuits, which is what the heuristic’s “can the harm be undone” question is asking about.

A verdict in any of the three columns is incomplete without its conditions. An adopt is only honest with them attached: what gets monitored after launch and who sees those numbers, who can override the system’s output, what result triggers rollback, and how the agency gets its data and workflow back if the tool is retired. A refuse carries the matching obligation to say what would change the answer, whether that is a finished governance policy, a different contract, or evidence from someone else’s deployment.

In many low-resource agencies the verdict was never on the table. Organizations receiving HUD Continuum of Care or Emergency Solutions Grants funds are required by statute to participate in HMIS, the Homeless Management Information System, as a condition of the funding. The 21st Century Cures Act required every state to implement electronic visit verification for Medicaid personal care services by January 2020, with escalating cuts to a state’s federal match for missing the deadline. EVV in practice often means a GPS-enabled app logging when and where a home-care visit happens, and Data and Society’s study of the rollout found that monitoring the worker also indirectly tracks the client’s movements, with a chilling effect on the daily lives of the disabled and older people receiving care.

A mandate takes the adopt/defer/refuse question away and leaves the conditions question intact. Even the HMIS requirement has negotiated exceptions: victim service providers are prohibited from entering survivor data into HMIS and must use a comparable database instead, and each Continuum of Care selects its own software. For a mandated system, the heuristic’s questions point at the terms: which product, what data beyond the required minimum, who sees it, and what the override and appeal paths look like.

The conditions question has a method behind it. Pawson and Tilley’s realist evaluation asks what works, for whom, in what circumstances, and why, and reads any outcome as a mechanism firing in a context; Week 9 takes the method up as an evaluation tool.

Vendors and procurement

Deployment usually means a vendor, and the contract decides more than the model does. Indiana signed a $1.3 billion contract with IBM to privatize and automate welfare eligibility; the redesign removed caseworker discretion, the state’s error rate tripled, mostly through wrongful terminations, during a recession-era application surge, and the arrangement ended with a judge awarding Indiana roughly $78 million after ruling IBM breached the contract.

Concentration compounds the risk. Deloitte runs Medicaid eligibility systems in 25 states covering 53 million enrollees, under contracts worth at least $5 billion, with failures reported in Kentucky, Rhode Island, and Colorado, where a 2023 audit found errors in 90 percent of sampled notices. After a federal judge ruled Tennessee’s Deloitte-built system had illegally cut Medicaid coverage, advocacy groups asked the FTC to investigate the vendor itself.

Vendors also leave. Woebot held an FDA Breakthrough Device designation and had raised a $90 million funding round, and it still retired its consumer app in June 2025, with the founder citing regulatory costs and faster-moving general-purpose models; about 1.5 million lifetime users lost the service. An agency that builds intake around a product it does not control inherits that possibility, which is what the heuristic’s exit question is for.

State law increasingly constrains what an agency may adopt in the first place. The Mental Health AI Policy Project’s 50-state tracker of AI mental-health legislation followed 44 bills as of July 2026, including 22 enacted laws, and it filters by state, so an agency weighing a chatbot or documentation tool can check what its own legislature already requires or prohibits.

The gap between task and deployment

The optimist RCTs measure a task under controlled conditions; deployment happens in an under-resourced agency with vulnerable clients, procurement contracts, and staff nobody trained. Goldkind, Ming, and Fink sort the sector’s AI talk into hype, harm, and hope, and argue that algorithmic harms surface disproportionately where AI systems meet pre-existing structural inequality, which describes most human-services caseloads. AI can compress skill gaps among supported users while access and governance gaps widen inequality across a population: six months of search data show generative-AI uptake clustered in coastal metros and higher-income, higher-education counties, with the South, Appalachia, and much of the Midwest lagging. A tool that helps its users while missing the people already worst served raises average performance and still widens the equity gap.

The divide runs between organizations too. In a 2025 TechSoup and Tapp Network benchmark survey of more than 1,300 nonprofit professionals, nonprofits with annual budgets over $1 million were adopting AI tools at about twice the rate of smaller organizations, 66 versus 34 percent, and 43 percent of the organizations surveyed relied on one or two staff members for their IT and AI decisions.

The boundary of what the model can do is itself part of the gap. Dell’Acqua and colleagues ran 758 consultants through tasks on both sides of what the tool could actually handle, and found large gains inside that boundary and worse performance than the unaided control outside it. The boundary is invisible from inside the conversation, since the model answers the same way on both sides of it.

Perception adds its own gap. An RCT with 16 experienced open-source developers on 246 real tasks found AI tools made them 19 percent slower on their own mature codebases; they had predicted a 24 percent speedup beforehand, and after finishing they still believed the AI had made them about 20 percent faster. Self-report is a poor instrument for measuring what these tools do, which matters because most adoption surveys are self-report.

Task level: controlled RCTs

Noy and Zhangwriting time fell 40 percent
Brynjolfssonnovice productivity rose 34 percent
Therabotsymptoms fell

Deployment: real agencies

MiDAS40,000 people falsely accused
TennCarecoverage cut illegally
Woebotshut down in 2025

The same tasks that succeed in a controlled trial meet governance, procurement, staff training, vulnerable clients, and a 12-month clock once an agency deploys them, and that gap is what separates the two rows.

Adoption as an implementation problem

Whether a technology helps rarely turns on the technology alone. NASSS asks why health technologies end in nonadoption or abandonment even when the tool works, tracing failure across seven domains that run from the condition itself out to the wider system, most of which no trial ever tests. EPIS, developed in child welfare, walks adoption through exploration, preparation, implementation, and sustainment; the motion’s 12-month clock covers only the first three of those phases. CFIR, the field’s consolidated framework, organizes 37 constructs across five domains for asking what works where and why.

Pahlka’s Recoding America makes the government-side version of the argument: American service-delivery failures live in implementation, in an industrial-era culture that separates the people who write policy from the people who build and run the systems that deliver it.

The equity frameworks ask a different question. The Health Equity Implementation Framework, first applied to hepatitis C treatment uptake among Black VA patients, found standard implementation barriers operating alongside equity-specific ones such as historical mistrust, and Baumann and Cabassa propose centering reach from the start: who an implementation reaches, who it burdens, and who it leaves out.

Reading evidence as an administrator

Citing a real study badly loses debates and misleads boards. The first check is the comparator. Therabot beat a waitlist, the weakest possible control, in its 210-person trial; the reductions were real (51 percent in depression symptoms, 31 percent in anxiety), and a waitlist comparison cannot say whether the chatbot beat a workbook, a weekly phone call, or a human therapist.

The second check is which tool was tested. Therabot was expert-built, fine-tuned, and clinician-monitored, and trial participants averaged about six hours of engagement, roughly eight therapy sessions’ worth; a generic chatbot deployed without that scaffolding borrows the trial’s credibility without its conditions.

The third check is whether the headline number is a measurement or a projection. The Nigeria tutoring pilot measured about 0.31 SD of learning gain in six weeks, 0.23 SD in English, at about $48 per student; the widely quoted “2.23 SD per year” is an extrapolation from that six-week result. The last check is whether a task-level gain says anything about population-level equity, which the adoption-divide evidence says it does not.

The evidence packets

Your team gets one of two evidence packets a week before class.

The affirmative stack opens with Noy and Zhang (Science, 2023), who ran an RCT with 453 professionals on mid-level writing tasks; ChatGPT cut time 40 percent, raised quality 18 percent, and narrowed the gap between stronger and weaker writers. Brynjolfsson, Li, and Raymond (QJE, 2025) tracked 5,179 customer-support agents through a real rollout; productivity rose 14 percent on average and about 34 percent for novices, with requests to escalate to a manager falling 25 percent, while the most experienced agents gained almost nothing. The World Bank’s Nigeria pilot produced measurable learning gains at about $48 per student, and Therabot showed significant symptom reduction in the first RCT of a generative therapy chatbot, with self-reported alliance comparable to human therapy.

Redesigning administrative technology has its own trial evidence. A randomized SNAP experiment by Giannella and colleagues found that letting applicants schedule flexible interviews, instead of waiting on one unscheduled call from a caseworker, raised approval rates by six percentage points, roughly doubled early approvals, and raised longer-term participation by more than two points. That result is Herd and Moynihan’s theory of administrative burden tested experimentally: burdens are deliberate policy choices that fall disproportionately on disadvantaged populations, a pattern the book traces through cases from voter ID to abortion access. Whether someone who qualifies for a benefit actually receives it depends on the learning, compliance, and psychological costs the process loads onto them, and intake technology can raise or lower any of the three.

The negative stack opens with Michigan’s MiDAS, which falsely accused roughly 40,000 residents of unemployment fraud, wrong about 93 percent of the time. For nearly two years the $44.4 million system made fraud determinations with no human review at all, quintupling fraud accusations while a 400 percent penalty structure pushed the agency’s fraud-penalty collections from about $3 million to over $69 million. Bridge Michigan followed individual claimants, including one man who had more than $14,000 garnished and was told he still owed $70,000, mostly penalties. Accountability took a decade: a $20 million settlement covering roughly 3,000 class members was approved in January 2024, after the Michigan Supreme Court first had to rule that harmed workers could sue the state at all.

TennCare Connect, a $400 million eligibility system, illegally cut Medicaid coverage for thousands, per a federal judge in 2024, and its builder made similar systems in more than 20 states; one Nashville mother with chronic anemia lost coverage because the system autofilled a wrong address, leaving her uninsured for nearly two months before a preeclampsia diagnosis. Los Angeles ranks unhoused people for scarce housing with a survey-based score, and among young adults in 2021, 67 percent of white respondents reached the highest-priority tier versus 56 percent of Latino and 46 percent of Black respondents, in a county where Black people are 30 percent of the homeless population and 9 percent of residents overall.

Woebot, the most evidence-backed mental-health chatbot on the market, shut down its app in June 2025 under regulatory costs, so strong evidence did not guarantee survival. One stack measures what the tool does to a task; the other measures what happens once that task-level gain meets an underfunded agency, a vendor contract, and a client population with little power to contest an error. A debate team that can name which of those conditions hold in a real agency, on a 12-month clock, beats a team that only cites the larger effect size.

Before class

Readings

Evidence packets

Your team, your side, and your evidence packet arrive a week ahead. Read the packet before lab; in-class time goes to argument.

In class

Seminar

Anonymous vote, then we build the adopt/defer/refuse heuristic on the board using the intake workflow, followed by the clinic on reading RCTs critically. Lauri Goldkind (Fordham), editor of the Journal of Technology in Human Services, then speaks on what AI adoption looks like inside human-services organizations, with time for your questions.

Activity

Debate Lab 1

Teams caucus with their packets, then run timed rounds: opening constructives, evidence presentations, cross-examination, rebuttals, and a closing synthesis in which each side concedes the other’s strongest point and argues it still wins on the 12-month timeline. A panel of three classmates judges against a rubric rewarding evidence fit, direct engagement, sociotechnical reasoning, and intellectual honesty (judges rotate across the two debate labs, so everyone debates once and judges once). The class then revotes, and we look at who moved.

Further reading