Module 2, Sep 30, Seminar and lab. Partner matches are announced this week.
AI literacy
Background
When a chatbot gets the same health question in English and in Arabic, the Arabic answer is measurably worse. In that study, bilingual experts scored four chatbots on 15 questions about five infectious diseases; English answers averaged 4.6 on a 5-point scale, in the excellent range, while Arabic averaged 4.1. The same drop has been measured for Spanish, on real patient questions about epidurals, where bilingual obstetric anesthesiologists rated Spanish answers significantly lower than English ones, and for Chinese and Hindi across three expert-annotated health question sets. The tool performs worst for many of the people social-impact organizations serve.
The evidence complicates that response. In a 2024 Nature Medicine study with 2,280 participants, labeling identical, accurate medical advice as AI-generated made people rate it less reliable and less empathetic and made them less willing to follow it, even when a physician was said to be supervising. Over-warning has victims too, and a community taught only fear forfeits real access.
Distrust is already the majority position. When Pew Research Center surveyed 5,410 US adults and 1,013 AI experts in 2024, 17 percent of the public expected AI to have a positive effect on the United States over the next 20 years, against 56 percent of the experts, and the share of adults more concerned than excited about AI in daily life reached 50 percent by mid-2025, up from 37 percent in 2021. Majorities of both the public and the experts told Pew they want more control over how AI is used in their lives.
The durable skill is calibration: knowing when an answer is probably fine, when it needs checking, and when the question belongs to a person. Specific prompt tricks decay quickly, because effects found for one model or one prompt phrasing often disappear on the next version. Calibrated judgment about when to trust an answer is what carries over instead.
Key ideas
Calibrated trust
Two findings anchor the practice. In Buçinca and colleagues’ study of 199 participants, designs that forced people to engage before seeing the AI’s answer, making your own guess first among them, cut overreliance on wrong answers, while attaching explanations to answers did not. Two details complicate the win: participants rated the most effective designs the least likable, and the benefit was larger for people who already enjoy effortful thinking, so a tool that helps on average can still leave the least-engaged users most exposed.
The second finding concerns expressed uncertainty. In a preregistered experiment with 404 participants answering medical questions through a fictional LLM search tool, a model saying “I’m not sure, but…” lowered trust and agreement while measurably improving accuracy; the same hedge phrased impersonally worked in the same direction, more weakly. Expressed uncertainty is useful information.
A 2025 review of automation bias across 35 studies reaches the same conclusion from the deployment side. Explanation features do not reliably reduce the bias, and explanations that are too technical, too demanding, or too simple can all reinforce misplaced trust, especially for less experienced users; engagement with the task is what keeps human judgment in the loop.
Offloading has a measurable cost of its own. In an EEG study of 54 people writing essays with an LLM, with a search engine, or unaided, the LLM group showed the weakest neural connectivity, reported the lowest sense of ownership over their essays, and struggled to quote their own writing back minutes later; the study is a preprint, and its authors call the pattern cognitive debt. In practice all of this means making your own guess first, treating the model’s answer as a draft, and verifying anything that matters.
Ask AI or ask a person
For client-facing work, a harm-reduction line puts logistics and navigation (find the office, get the gist of a form, draft a question for the caseworker) in reasonable AI territory, while diagnosis, dosage, and immigration-legal exposure go to a qualified person. A crisis always goes to a human.
The crisis rule has current evidence behind it. In November 2025 the American Psychological Association issued a health advisory warning that generative chatbots and wellness apps are unpredictable in crisis situations; a companion test found that none of two dozen chatbots gave an adequate response to a user expressing suicide risk. The advisory calls for FDA oversight and a ban on AI impersonating licensed clinicians. The language gap raises the stakes of the triage line too, because AI translation does not satisfy the legal right to a qualified interpreter for anything clinical or legal.

The language gap
The gap is widest exactly where machine translation gets used most casually. When bilingual community evaluators rated 400 machine-translated emergency discharge instructions across seven languages, accuracy ran from 94 percent in Spanish down to 82.5 percent in Korean, 67.5 percent in Farsi, and 55 percent in Armenian. The failures were not subtle: “ibuprofen” came out as “anti-tank missile” in Armenian, and “Coumadin,” a blood thinner, came out as “soybean” in Chinese.
A 2026 analysis of AI-mediated medical interpreting lays out why the drop matters clinically. Performance is worst on low-resource and Indigenous languages, idioms get flattened (an idiomatic “chest tightness” rendered as generic discomfort, a phrase a clinician needs precisely), and routing client speech through third-party servers creates HIPAA exposure. The analysis recommends AI only as a supplement when a human interpreter is unavailable, and argues that normalizing a lower standard of care for non-English speakers is itself an equity harm.
For the triage sort, this puts a first-pass translation for gist in verify-first territory and anything clinical or legal with a qualified interpreter. The client who needs the translation is usually the one least positioned to catch its errors, which is why the verification burden belongs to the worker.
Error, accountability, and participation
Every algorithm that sorts people makes false positives (flagging someone who should have passed) and false negatives (passing someone who needed the flag). Pushing one rate down pushes the other up, so choosing the balance point is a values decision about who absorbs which harm. People’s preferences over that tradeoff differ measurably: in large experiments in the US and Norway, most people weighted wrongly denying a deserving person more heavily than wrongly approving an undeserving one, with the balance shifting by country and political orientation.
The COMPAS recidivism score shows the tradeoff in the record. ProPublica’s analysis of more than 7,000 people arrested in Broward County found Black defendants falsely labeled future criminals at 44.9 percent versus 23.5 percent for white defendants, while white defendants who did reoffend had been labeled low risk at 47.7 versus 28.0 percent; overall accuracy was 61 percent. Models learn from recorded history, so when no social worker or affected community member sat in the room where thresholds were set, the model reproduces the record.
A caseworker’s error has an accountable author, with supervision, appeal, and a license behind it; an algorithmic error often answers to no one. Michigan’s MiDAS system issued roughly 40,000 false fraud accusations, wrong about 93 percent of the time, while the state cut its human fraud-review staff by about a third, and at least 11,000 affected families filed for bankruptcy before anyone could be held to account.
Private systems raise the same question of who answers for an error. A class action alleges UnitedHealth’s nH Predict algorithm, which predicts post-acute care length of stay from a database of 6 million patients, has a 90 percent error rate on the denials that get appealed, while only about 0.2 percent of policyholders ever appeal at all. In Allegheny County’s child-welfare screening tool, an AP investigation found the algorithm flagged 32.5 percent of Black children for mandatory investigation versus 20.8 percent of white children, and caseworkers disagreed with its scores about a third of the time. Week 3 takes the fairness tradeoffs apart in detail.
Misinformation and teach-back
AI fakes surface polish, so “it looks off” no longer works as a detector. In a June 2025 Pew survey of 5,023 US adults, 76 percent said it is extremely or very important to be able to tell whether pictures, videos, and text were made by AI or by people, and 53 percent said they were not confident they could tell the difference.
The practice with the most trial evidence is prebunking: teaching people to recognize manipulation techniques, such as emotional language, false dichotomies, and scapegoating, before they meet them. Seven preregistered studies with about 30,000 total participants tested short prebunking videos, and a YouTube ad campaign reaching about 5.4 million users improved technique recognition by 5 percent at roughly five cents per view. The approach held up in a field experiment across 12 EU countries with 19,735 participants, run around a real campaign that reached more than 120 million YouTube viewers before the 2024 EU elections; even 20-second versions of the videos worked.
The effect sizes are on the record in both settings. In the lab studies, the videos improved recognition of manipulation techniques with effects between 0.28 and 0.68 standard deviations; in the EU field deployment, which targeted scapegoating, decontextualization, and discrediting, the gains were smaller, roughly 0.08 to 0.38, but held across all thirteen surveys and improved sharing decisions as well as discernment, with the 50-second videos working more consistently than the 20-second ones. Per person the effects are modest; the method’s advantage is that they held up when delivered as ordinary video ads at population scale.
The second practice is lateral reading, leaving the page to check who else says so. Wineburg and McGrew found that professional fact checkers evaluate an unfamiliar site by opening new tabs and reading less of any single source while triangulating across many, while students and even historians stayed on the page in front of them. The routine teaches quickly: a one-hour self-directed module moved older adults’ headline-discernment accuracy from 64 to 85 percent during the 2020 election, while a control group moved from 55 to 57.
Language shapes exposure here too. Factchequeado, a fact-checking nonprofit launched to close the Spanish-language misinformation gap for the more than 68 million US Latinos, partners with over 140 outlets across 27 states and Puerto Rico and runs a WhatsApp chatbot; immigration is a top misinformation target.
The communication half is teach-back, asking your listener to say the idea back in their own words so misunderstandings get caught in the room. A systematic review of 20 studies found positive results in 19 of them, across comprehension, medication adherence, and reduced hospital readmissions.
AI in social work practice
The profession is using these tools ahead of its own guidance. A national survey of 1,179 social workers, fielded October 2025 to February 2026 by UT Austin’s Moritz Center with NASW, found most already using AI in their current role, largely for correspondence, reports, and documentation, and about two-thirds named clear ethical-use guidelines as the profession’s most pressing AI-related need.
Educators report a similar pattern. Six social work educators who documented their ChatGPT use across a semester rated it useful in 85 percent of interactions, most often for teaching support. Rodriguez and colleagues have proposed adding generative AI competence as a tenth core competency to CSWE’s 2029 accreditation standards, and Ahn and colleagues argue social workers need AI literacy whether or not they ever use AI directly, because these systems are reshaping the conditions their clients live under.
Before class
Readings
- Campos, H., & Salmi, L. (2025). Critical AI health literacy as liberation technology. NAM Perspectives.
- Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think. Proceedings of the ACM on Human-Computer Interaction.
- Sallam, M., et al. (2024). Language discrepancies in the performance of generative artificial intelligence models. BMC Infectious Diseases.
In class
Seminar
Vote, then work through the six moves with their evidence: a prebunked wrong answer with the mechanism revealed, the calibration studies, the ask-AI-or-ask-a-person line, the language gap, lateral reading, and teach-back. Midway, the room splits and argues the proposition from assigned sides.
Lab: practice and governance
Four reps run in sequence.
- Answer first. Answer a realistic benefits question yourself before asking an LLM, then compare.
- The triage sort. Your group of three gets a deck of about fifteen client-facing scenario cards (find the county office, which documents for a SNAP renewal, is this rash dangerous, will using this benefit affect my green card, translate this consent form for signing) and sorts each into one of three lanes, AI-OK, verify-first, or human-only, defending every boundary; the trap cards get revealed at the end, and they sit exactly where AI performance is worst and the client is most exposed. A solo version to practice on is online, the Ask AI or ask a person sorter.
- Draft the AI-disclosure form. The class co-drafts its AI-disclosure form (what counts as reportable use, the required fields, the clause to verify and own every output), adopted provisionally for the Case Brief.
- Teach-back and revote. A short teach-back role-play and the revote close it out.
Lab: prompt writing and verification practice
Prompt engineering is the field’s name for wording a request to get a better result, and the name oversells the practice, since what looks like engineering is closer to structured trial and error, and a trick that works on one model version often fails on the next (Meincke et al., 2025). The durable skill underneath is understanding why wording changes a model’s output and verifying what comes back. Working from the before-class prompting tutorial as a warm-up, each pair takes one realistic human-services writing task and runs the same prompt through four revisions: add a role and audience, add context and constraints, add a worked example, and add an explicit ask to flag anything unverified, noting what changed after each step. Nothing in the final draft counts as done until a pair member fact-checks it against a real source, the disclose and verify rule that governs every AI-assisted deliverable in the course.
Further reading
- Kim, S. S. Y., et al. (2024). “I’m not sure, but…” FAccT ‘24.
- Romeo, G., & Conti, D. (2025). Exploring automation bias in human-AI collaboration. AI & Society.
- Reis, Reis & Kunde (2024). Influence of believed AI involvement on the perception of digital medical advice. Nature Medicine.
- Kosmyna, N., et al. (2025). Your brain on ChatGPT: Accumulation of cognitive debt. arXiv preprint.
- Jin et al. (2024). Better to ask in English. ACM Web Conference ‘24.
- Taira, B. R., Kreger, V., Orue, A., & Diamond, L. C. (2021). A pragmatic assessment of Google Translate for emergency department instructions. Journal of General Internal Medicine.
- Lopez Vera, A. (2026). Ethical risks and structural implications of AI-mediated medical interpreting. JMIR AI.
- Angwin, Larson, Mattu & Kirchner (2016). Machine Bias. ProPublica.
- Angwin (2022). The Seven-Year Struggle to Hold an Out-of-Control Algorithm to Account. The Markup. This piece becomes required reading in Week 3.
- Ho, S., & Burke, G. (2022). How an algorithm that screens for child neglect could harden racial disparities. Associated Press, via PBS NewsHour.
- CBS News. (2023). UnitedHealth uses faulty AI to deny elderly patients medically necessary coverage, lawsuit claims.
- Roozenbeek, J., et al. (2022). Psychological inoculation improves resilience against misinformation on social media. Science Advances.
- Communications Psychology. (2026). Video inoculation against election misinformation across 12 EU nations.
- Wineburg, S., & McGrew, S. (2019). Lateral reading and the nature of expertise. Teachers College Record.
- Moore & Hancock (2022). A digital media literacy intervention for older adults improves resilience to fake news. Scientific Reports.
- CrashCourse. (2019). Check yourself with lateral reading (video, 14 min). Episode 3 of the Navigating Digital Information series, made with MediaWise, the Poynter Institute, and the Stanford History Education Group.
- Factchequeado. About Factchequeado.
- Talevski, J., et al. (2020). Teach-back: A systematic review of implementation and impacts. PLOS ONE.
- Jena, B. (2022). Would you rather see a computer or a doctor? (Freakonomics, M.D., 32 min).
- American Psychological Association. (2025). Artificial intelligence, wellness apps alone cannot solve mental health crisis (health advisory).
- Ahn, E., Choi, M., Fowler, P., & Song, I. H. (2025). Artificial intelligence (AI) literacy for social work: Implications for core competencies. JSSWR.
- Creswell Báez, A., et al. (2025). Social work educators innovating with generative AI. Journal of Social Work Education.
- NASW-Illinois. (2026). National survey finds most social workers already using artificial intelligence.
- Mollick, E. (2023). Working with AI: Two paths to prompting, paired with Meincke, L., et al. (2025). Prompt engineering is complicated and contingent.
- Zamfirescu-Pereira, J. D., Wong, R., Hartmann, B., & Yang, Q. (2023). Why Johnny can’t prompt: How non-AI experts try (and fail) to design LLM prompts. CHI ‘23.
- Anthropic. Prompt engineering overview. This is the warm-up tutorial for the second lab.
- OpenAI. Prompt engineering guide.