Module 2, Sep 23, Seminar and lab.

Language model mechanics

Background

Bender and colleagues named these systems stochastic parrots, which stitch together linguistic form with no communicative intent and no grounding, and mistaking that fluency for understanding causes real harm. Their 2021 paper also priced the race toward ever-larger models, citing an estimate that training one large neural language model emits roughly 284 tons of CO2 and that a 0.1-point gain on a translation benchmark can cost about $150,000 in compute, alongside the documentation debt of training on web text nobody has curated or read.

Bender has since put boundaries on the claim. In an interview five years after the paper, she limits the term to text-generating language models (chess engines and AlphaFold were never the target) and restates its core: when the output makes sense, it is because a reader is making sense of it.

The counterevidence comes from interpretability research. Anthropic has traced internal structure that looks like planning several words ahead, with the model activating features for candidate rhyming words before it finishes a poetry line, and one shared concept space serving English, French, and Chinese. The full set of case studies documents multi-hop reasoning, with Texas appearing as an unstated intermediate step between “the state containing Dallas” and its capital, though the authors report their method worked on only about a quarter of the prompts they tried. The same research found cases where the model’s stated chain of reasoning did not match what it actually computed, so even a model that shows its work can misreport it.

Other evidence lands in between. In one Harvard and MIT study, a model produced accurate Manhattan driving directions without ever forming a coherent street map, and its performance collapsed once 1 percent of streets were blocked. Some researchers respond by redesigning the tool around its limits; one 2025 position paper proposes “reasonable parrots”, language models built on argumentation-theory principles so they exercise a user’s critical thinking instead of substituting for it.

As of mid-2026, courts had logged roughly 1,762 filings containing citations to cases that do not exist, invented by AI tools and submitted by lawyers who never checked. Social-impact work also goes on the record, in court reports, benefits appeals, and grant budgets.

AI as normal technology

Perspectives on AI sit at very different levels. At one end are forecasts of a superintelligence that becomes a separate kind of agent, a framing that runs from Bostrom’s account of an intelligence explosion to the 2023 statement, signed by many of the field’s leaders, that mitigating the risk of extinction from AI should be a global priority. A second register reads that fear itself as the object of study: Blix and Glimmer argue that the AI nightmare says more about capitalism than about the technology, an “AI” imagined as an autonomous agent standing in for capital, and that the doom talk of the managerial class tracks its worry that automation is now reaching the judgment and oversight work that legitimates its position. A third register treats the framing itself as the problem: Arvind Narayanan and Sayash Kapoor argue for treating AI as normal technology: “To view AI as normal is not to understate its impact,” since electricity and the internet were transformative and remain normal, but to set aside the habit, shared by the utopian and dystopian readings, of treating AI “akin to a separate species, a highly autonomous, potentially superintelligent entity.” Their claim is that impact arrives through diffusion, paced by the slow adaptation of people, organizations, and institutions, so those stay the main point of control. The course works in that register, following Week 1’s point that a technology passes from frontier to infrastructure; the telephone and the computer were once emerging systems whose consequences were hard to read. The course asks of AI what it asks of any technology: who builds it, whose problem it is set to solve, whose data and labor it runs on, and who carries the cost when it fails.

A montage of UK newspaper front pages covering the July 2024 CrowdStrike global IT outage with dramatic headlines
UK newspaper front pages on July 20, 2024, after a faulty CrowdStrike software update disrupted airports, hospitals, and banks worldwide. An ordinary software failure was framed as “the day the world stood still.” Front pages of various UK newspapers, July 2024.

The comparison to electricity has a literature behind it. Economists call these general-purpose technologies, and Bresnahan and Trajtenberg named the pattern in 1995: a technology spreads across most sectors, keeps improving, and makes other people’s innovations possible. The pattern includes a long delay before any of it shows up as output. Paul David’s account of electrification shows that factories gained little from electric motors while they kept the layout built around a central steam shaft, and the gains arrived only once the buildings, the machines, and the work were rearranged around what the new motors allowed, which took decades. Solow described the same lag in 1987 in a line about seeing the computer age everywhere except in the productivity statistics.

Normal in this argument is not a claim about speed. ChatGPT reached about 100 million monthly users within two months of launch, which UBS analysts called the fastest ramp they could recall in a consumer internet app; TikTok took about nine months to that mark and Instagram two and a half years. Bick, Blandin, and Deming ran the first nationally representative US survey of generative AI use and measured each technology two to three years after its first mass-market product. Generative AI stood at 45.5 percent overall use against 19.7 percent for personal computers in 1984, and they conclude that “work adoption of generative AI has been as fast as the personal computer (PC), and overall adoption has been faster than either PCs or the internet.”

At work the rates are close, 32.1 percent for generative AI against 25.1 percent for PCs. At home they are 37.7 percent against 5.5 percent. People are taking up generative AI in their own lives well ahead of their employers, and the PC numbers in the same survey run the other way.

Whether language models belong on that list is an open question with numbers on both sides. Eloundou and colleagues estimate in Science that about 80 percent of US workers have at least 10 percent of their work tasks exposed to language models, and about 19 percent could see half their tasks affected. Acemoglu’s macroeconomic estimate puts the resulting gain in US total factor productivity at no more than 0.66 percent over ten years, and revises that down to less than 0.53 percent, while Goldman Sachs projects a 7 percent rise in global GDP and a 1.5 point lift in productivity growth over a decade. Two credible estimates of the same technology sit an order of magnitude apart. Acemoglu and Johnson’s Power and Progress makes the further argument that who gets the gains has never been decided by the technology, and that the broad prosperity of earlier transitions came from the institutions and countervailing power built around them.

Whether the reorganization has started showing up in employment is contested. Brynjolfsson, Chandar, and Chen report a 16 percent relative decline in employment for workers aged 22 to 25 in the most AI-exposed occupations since late 2022, using payroll microdata, with growth for young workers in occupations where the tools complement the job. The authors themselves have since weighed interest rates and the general hiring slowdown as alternative explanations, so the finding is suggestive and not settled. Week 6 takes up the trial evidence on what these tools do to a task, and what happens to it in a real organization.

Treating AI as normal technology does not make it benign, and the critical scholarship on AI hype is another of these levels the course reads closely. Karen Hao’s Empire of AI traces how the industry concentrates power, compute, and capital in a few firms, and reports the costs carried elsewhere, from the low-paid workers who label and moderate training data to the water and energy that data centers draw. A separate question is what the systems actually do. Whether they understand what they produce is genuinely open: Mitchell and Krakauer survey the debate and argue that understanding is a matter of degree and kind rather than a yes-or-no verdict the current tools can settle, and this week’s background sets the stochastic-parrots critique against interpretability evidence of planning and multi-step reasoning. What the course guards against is the salvation view that AI belongs in every setting, and the opposite insistence that the profession should not engage it at all. A social worker meets these tools already in the case file, the benefits portal, and the courtroom, and the working task is to judge a specific system in a specific setting, which is where Narayanan and Kapoor’s AI Snake Oil offers a method, telling what AI can do apart from what gets sold in its name. The course does not ask students to endorse or reject AI as such. Its aim is that they understand these systems well enough to form their own judgment about whether and where such tools belong in social services.

The profession’s own rules already say something about this, though less than is sometimes claimed for them. The 2017 Standards for Technology in Social Work Practice require social workers who use technology to provide services to “obtain and maintain the knowledge and skills required to do so in a safe, competent, and ethical manner,” and the NASW Code of Ethics at 1.04(d) carries the same obligation. Both were written before ChatGPT existed, and both attach to the social worker who uses a tool, so neither settles what the profession as a whole should do about AI.

Elisa Borah and colleagues at UT Austin’s Moritz Center, with NASW, the Department of Veterans Affairs, and the Department of Defense, surveyed 1,179 US social workers between October 2025 and February 2026. They found 63.5 percent using AI in their current role, many of them several times a week or daily, across every practice area in the sample, most often for correspondence and reporting, clinical documentation, administrative support, data analysis, and research. Asked what would help, 66.8 percent chose clear guidelines on the ethical use of AI, the most endorsed item on the list, 55.5 percent asked for better privacy and security measures for client data, and respondents named privacy, bias, transparency, client safety, and overreliance on automated decisions as their main concerns. The report describes a readiness gap, with use already high and many practitioners feeling underprepared to use these tools ethically, and quotes one of them: “We are in desperate need of guidance on ethical use.” Rodriguez, Goldkind, Victor, Hiltz, and Perron have proposed adding a generative AI competency to CSWE’s accreditation standards for 2029.

The evidence that understanding the mechanism improves anyone’s judgment is thin, and what exists cuts against it. Green reviewed 41 policies requiring human oversight of government algorithms and found that the people assigned to oversee were unable to perform the function, so the requirements largely legitimated the systems they were meant to check. Elish documents a related pattern she calls the moral crumple zone, in which responsibility for a failure settles on the human closest to an automated system, the one with the least control over what it did. Neither finding settles whether this week’s material helps. The case for AI literacy in social work currently rests on the professional standards and on how many practitioners already use the tools; nobody has yet shown that the literacy itself changes what practitioners do.

Key ideas

These are the parts the seminar and the lab use. Language models in depth carries the rest: where deep learning came from, what scale bought, what retrieval and agents add on top, and the argument over whether any of this resembles human intelligence. A generative language model is one family of AI among several; Kinds of AI systems sets it beside the predictive risk models and the NLP and vision systems the word “AI” also covers, and explains why ChatGPT came to stand in for all of them.

Next-token prediction

Text gets chopped into tokens before the model touches it. The standard method, byte-pair encoding, builds the vocabulary by repeatedly merging frequent character pairs until it reaches a target size of roughly 50,000 tokens, which is why “dehumidifier” splits into four subword pieces. Letters are largely invisible to the model, which is why it can spell a word correctly and still miscount letters: the same article has GPT-4 counting a run of 29 letters as 30, and the widely repeated case is a model miscounting the r’s in “strawberry.”

Each token then becomes a long list of numbers positioned so related things sit near each other. In GPT-3’s largest version, every word lives in a 12,288-dimensional vector space across 175 billion parameters, and meaning becomes geometry: “biggest” minus “big” plus “small” lands near “smallest.” The core loop of every chatbot is the same: produce a probability for every possible next token, pick one, append it, repeat.

Attention, the transformer’s key move, lets each token consult the others, and a temperature setting controls how adventurously the model samples, which is why one prompt can produce three different answers. All of it is learned first from a huge corpus (GPT-3’s training text is equivalent to more than 2,600 years of continuous human reading) and then from human labelers who shape tone and helpfulness, so whose text and whose labels matter at every step. This becomes the bias mechanism in Week 3.

Textwhat you typed
Tokenschopped into pieces
Vectorseach piece becomes numbers
Attentioneach token consults the others
Probabilitiesa score for every possible next token
Sample onepick a token from the scores

The sampled token is appended to the tokens at step 02 and the whole loop runs again, one token at a time, for as long as the answer keeps going.

Diagram of the transformer architecture with attention blocks
The transformer architecture, with its stacked attention blocks, underlies today’s large language models. In lab you will watch a live one pick its next token. Diagram: dvgodoy (dl-visuals), via Wikimedia Commons, CC BY 4.0.

The transformer

The architecture every current chatbot runs on was published by Google researchers in 2017 under the title Attention Is All You Need. The models it replaced read a sentence one word at a time, each step waiting on the last, and the paper names the cost of that design directly: “This inherently sequential nature precludes parallelization within training examples.” Dropping recurrence let a whole sequence be processed at once across many GPUs, which is what made training at the current scale possible.

Attention itself is a matching operation. Each token sends out a query describing what it is looking for, every token offers a key describing what it has, and the strength of the match between a query and a key decides how much of that token’s value gets mixed in. The 2017 base model ran eight of these attention heads side by side in each layer, so a token could track several kinds of relation at once, and because nothing in the mechanism knows about word order, position has to be injected into the input as its own separate signal. Most of the parameters live in the plain feed-forward layers between the attention blocks, and Geva and colleagues found those layers behave like key-value memories, where each key matches a readable pattern in the training text and its value tilts the output distribution.

The original models were small by current standards. The base model held about 65 million parameters and trained in 12 hours on eight P100 GPUs; the larger one held 213 million and trained for three and a half days, reaching 41.8 BLEU on English-to-French translation. Machine translation was the whole task in that paper, which claims nothing about understanding or conversation.

Every token attends to every other token, so the work grows with the square of the sequence length and doubling the context roughly quadruples the compute. Context windows grew anyway, from GPT-3’s 2,048 tokens in 2020 to Gemini 1.5 Pro’s 128,000 in February 2024, with up to 1 million offered to a limited group of developers and Llama 4 Scout’s 10 million in April 2025, which is why a model can now hold an entire case file and why holding it is expensive.

Hallucination and the training objective

The model’s job is likely text, and a fluent fabrication often scores as more likely than an honest “I’m not sure.” Ted Chiang’s description of the mechanism is a blurry JPEG of the web: the model keeps the general patterns of everything it read, drops the exact details, and renders the blur as fluent sentences. His essay opens with a 2013 case in which Xerox photocopiers using lossy compression silently altered room dimensions on scanned architectural drawings while the copies looked exact.

A 2025 OpenAI paper adds the training-side reason. Models are scored, in training and on benchmarks, the way test-takers are scored when blank answers earn zero, so confident guessing beats admitted uncertainty. The authors’ proposed fix is changing how benchmarks score, since a model currently loses points for saying “I don’t know.”

A 2025 letter in a medical journal, arguing about what these tools mean for medical writing, cites the benchmark rates from OpenAI’s own system card: hallucination fell from 12.9 to 9.6 percent across model generations, a reasoning-mode variant reached 4.5 percent, and the newest model hallucinated on 47 percent of fact-seeking tasks when it could not reach the web. The mechanism explains why the number does not reach zero.

Strengths of the mechanism

The same mechanism that fabricates citations is strong at tasks where form is the job: drafting a plain-language version of a dense notice, producing a first-pass translation for a human to check, and role-playing a hard conversation on demand.

The Trevor Project trains crisis counselors against a simulated teenager built on this mechanism. The first persona, Riley, a young person struggling to come out as genderqueer, was built from six months of research and thousands of role-play transcripts, with $2.7 million from Google.org, toward a goal of growing the volunteer corps from 700 to 7,000 digital crisis counselors, nearly 70 percent of whom serve nights and weekends. Fabrication carries little risk here because the simulated teen has no facts to get wrong.

Be My Eyes uses the mechanism to describe images for blind users, a population of more than 250 million people worldwide who once waited for a sighted volunteer to pick up. Unlike earlier object-recognition tools, it can hold a conversation about context, such as whether noodle ingredients look right or whether an object on the floor is a tripping hazard. Interviews with 26 blind users document where these tools fail (complex document layouts, non-English languages, cultural artifacts) and the verification workarounds users build, from cross-checking devices to enlisting sighted assistants; the authors argue AI tools should ship with explicit ways to contest errors.

Time and attention

Screen time changed daily life without anyone deciding it should. Common Sense Media’s 2021 census puts 8- to 12-year-olds at about five and a half hours of screen media a day and 13- to 18-year-olds at about eight and a half, and that share of childhood was never decided anywhere, it accumulated through millions of separate small choices.

Chatbot use has grown faster than that. ChatGPT reached 800 million weekly users by October 2025. An OpenAI study with economists at NBER reports more than 2.5 billion messages a day by mid-2025, roughly 10 percent of the world’s adults, and finds that non-work messages grew from 53 percent of usage to more than 70 percent in a single year. The categories are asking for information at about 49 percent of messages, getting a task done at about 40 percent, and expressing something personal at about 11 percent.

Teenagers use these systems for companionship at rates that are already measured. Common Sense Media surveyed 1,060 US teenagers and found 72 percent had used an AI companion and 52 percent used one regularly, with about a third finding those conversations as satisfying as talking to a friend and about a third having taken something serious to a chatbot instead of a person. Pew’s separate survey of 1,458 teenagers puts chatbot use at 64 percent, with 16 percent using one for casual conversation and 12 percent for emotional support, and the two surveys define the behavior differently enough that they should not be read as one trend.

Nobody can yet say what this does to people. A four-week study of about 1,000 participants run by OpenAI with the MIT Media Lab found heavier users reported more loneliness, more emotional dependence, and less socializing with people. The randomized conditions, voice against text and personal against neutral topics, produced no such differences; the association is with how much people chose to use it, so the study does not show that chatbot use causes loneliness. Sherry Turkle has argued for fifteen years that we come to expect more from technology and less from each other, and that the friction of real relationships is what people practice on.

Regulators have moved faster than the evidence. The FTC opened a 6(b) inquiry in September 2025 into how Alphabet, Character Technologies, Instagram, Meta, OpenAI, Snap, and xAI test companion chatbots for effects on children. California’s SB 243, signed in October 2025, with the provisions here effective January 2026 and a separate reporting mandate from July 2027, requires operators to disclose that a companion chatbot is not human, to maintain published protocols for suicide and self-harm content, and creates a private right of action. Garcia v. Character Technologies sits behind those statutes, brought by the mother of 14-year-old Sewell Setzer III after his 2024 suicide; in May 2025 the court rejected the argument that chatbot output is protected speech and let claims for product liability, negligence, and wrongful death proceed.

The fabricated citations case

By mid-2026, Damien Charlotin’s public database of AI hallucination cases had logged roughly 1,762 court decisions worldwide dealing with fabricated AI citations, accumulating since the first cases in 2023. Coverage runs from US federal and state courts to Canada, the UK, Spain, India, and Australia; in one Indian matter the database tracks, a lower court’s ruling was set aside for relying on case law that never existed. In case after case, a lawyer under deadline asked a chatbot for supporting cases, received perfectly formatted citations with plausible names and page numbers, filed them, and a judge discovered none of them existed. Sanctions follow, and the database logs each one.

The pattern has already reached benefits work. In Mavy v. Commissioner of the Social Security Administration, a brief filed for a disability appellant cited 19 cases, and the court found 12 of them fabricated, misleading, or unsupported. The judge revoked the attorney’s permission to practice before that court, struck the brief, and ordered apology letters to the judges whose names appeared on the invented opinions, while the client waited on a disability decision.

Courtrooms are just where the failure gets documented. In October 2025, KPMG published a report on agentic AI, and a citation audit found only 5 of its 45 citations accurately pointed to real sources; organizations named in the report, including the NHS and Transport for London, disputed claims attributed to them, and KPMG pulled it.

In this field, the equivalent failure would be a benefits appeal citing a regulation that was never written, a grant report carrying an invented statistic, or a resource list sending a client to a program that closed years ago.

Before class

Readings

In class

Seminar

A cold open puts half a sentence on the projector, and the room guesses the next word before the model reveals its ranked choices. The class then covers the mechanism in plain language (no math on the slides), why hallucination follows from it, and the parrot debate, with Bender’s critique and the interpretability evidence given equal footing. A short guest segment may be added (TBC).

Lab: drive the model

Using Transformer Explainer, which runs a real model live in your browser, you read off ranked next-token predictions, rerun one prompt at three temperatures, and watch attention resolve what “it” refers to in a sentence. The lab closes with a hallucination hunt. In pairs, you deliberately fish a full chatbot for confident fabrications (a real but not-famous researcher’s publication list, facts about a small local nonprofit) and verify every claim against a real source.

Screenshot of Transformer Explainer mid-interaction, showing attention flowing between tokens and ranked next-token probabilities
Transformer Explainer mid-run: attention weights connecting the input tokens, and the ranked probabilities for the next word, the live view lab uses. Screenshot of Transformer Explainer, July 2026.
Further reading