Second Nature. sociology in the age of AI — a working guide

Section 03 · The studies

Research that changed what we know

Twenty-seven studies, in order. Each is set out the same way — the question, the method, the finding, and the afterlife — because a finding is only as good as what happened when other people checked it.

Sociology's reputation rests on a surprisingly small number of studies that almost everyone has heard of. Some have held up remarkably well. Some turned out to be weaker than their fame. A few were, on close inspection, closer to theatre than science. Knowing which is which is part of thinking sociologically, and it matters more now that AI systems are trained on, and cite, the same popular summaries.

Each entry carries a status, a judgement based on the replication and critical literature as of 2026:

Scorecard — jump to any entry
StudyYearMethodStatus
Durkheim, Suicide1897Official statisticsQualified
Du Bois, The Philadelphia Negro1899Door-to-door survey, mixed methodsHolds up
The Hawthorne studies1924–32Workplace experimentsDoubtful (the “effect”)
The unemployed of Marienthal1933Community studyHolds up
The People's Choice1944Panel surveyQualified
Goffman, Asylums1961EthnographyHolds up
Milgram's obedience studies1963Laboratory experimentQualified
The Coleman Report1966National surveyQualified
Stanford Prison Experiment1971SimulationDoubtful
The strength of weak ties1973Survey, network theoryHolds up
“On Being Sane in Insane Places”1973Covert field experimentDoubtful
Bourdieu, Distinction1979Survey, correspondence analysisQualified
Hochschild, The Managed Heart1983Interviews, observationHolds up
Moving to Opportunity1994–Randomised experimentHolds up
The mark of a criminal record & name audits2003–04Field audit experimentsHolds up
MusicLab2006Online experimentHolds up
Obesity spreads through networks2007Longitudinal network dataContested
The 61-million-person experiment2012Platform experimentHolds up
Facebook emotional contagion2014Platform experimentQualified (tiny effect; ethics)
“Machine Bias” and COMPAS2016Data journalism, auditContested (and instructive)
Desmond, Evicted2016Ethnography + surveyHolds up
Gender Shades2018Algorithm auditHolds up
Bias in a hospital algorithm2019Algorithm auditHolds up
The Fragile Families Challenge2020Mass prediction challengeHolds up
Social capital and mobility2022Big data, 72 million usersHolds up (correlational)
The feed and polarisation2018–23Field and platform experimentsContested
Generative AI at work2023–25Field and lab experimentsQualified (early)

Durkheim, Suicide

Émile Durkheim · 1897 · France and Europe · official statistics

Question
If suicide is the most individual of acts, why are national suicide rates so stable from year to year, and why do they differ so consistently between groups?
Method
Comparison of official mortality statistics from several European countries and regions, by religion, marital status, family size, military service, and economic conditions. Durkheim systematically ruled out explanations based on climate, heredity, imitation and mental illness.
Finding
Suicide rates vary with the degree of social integration and regulation. Too little integration produces egoistic suicide (higher among Protestants than Catholics, among the unmarried than the married). Too much produces altruistic suicide (higher among soldiers). Too little regulation produces anomic suicide, which rises in economic crises and, strikingly, in sudden booms. Too much produces fatalistic suicide, mentioned only in a footnote.
Why it mattered
It was a demonstration that sociology could explain something no other science claimed. It established the use of quantitative comparison and the idea that rates are social facts.
Afterlife
Critics, notably Jack Douglas in The Social Meanings of Suicide (1967), argued that official statistics are themselves social products: a death recorded as suicide in one community may be recorded as an accident in another, especially where suicide carries religious stigma. Studies in the 1990s suggested some of the Protestant-Catholic gap was due to how deaths were classified. On the other hand, a 2018 study of nineteenth-century Prussia by Sascha Becker and Ludger Woessmann found that Protestantism did raise suicide rates even after accounting for such problems.

Qualified The specific numbers are shaky; the core idea that social integration protects against suicide is well supported and underpins modern public health research, including work on “deaths of despair”.

Du Bois, The Philadelphia Negro

W. E. B. Du Bois · 1899 · Philadelphia, Seventh Ward · household survey, history, mapping

Question
Commissioned partly by reformers who assumed Black residents were the cause of the ward's problems: what were the actual conditions of Black life in the city, and what produced them?
Method
Du Bois personally went door to door over roughly fifteen months in 1896–97, interviewing around 2,500 households with detailed schedules on family, work, income, housing, education and church membership. He combined this with historical research, census data, and hand-drawn maps classifying every household by social class.
Finding
Problems commonly attributed to Black character — poverty, crime, poor health — were the products of history (slavery and its aftermath), migration, and above all discrimination in employment, which shut even skilled Black workers out of decent jobs. Du Bois also showed the internal class differentiation of the Black community, which white observers did not see.
Why it mattered
It was one of the first sociological studies in the US to combine systematic empirical methods at this scale, and it made a structural rather than racial explanation of inequality.
Afterlife
Largely ignored by white academic sociology for decades. Now widely recognised as a founding work of American empirical sociology and of urban sociology.

Holds up Its method is still a model, and its argument about discrimination in labour markets is the ancestor of every modern audit study, including the algorithm audits below.

The Hawthorne studies

Western Electric; Elton Mayo, Fritz Roethlisberger, William Dickson · 1924–32 · Cicero, Illinois · workplace experiments

Question
Initially: does better lighting raise productivity? Later: what makes workers more productive?
Method
A series of studies at Western Electric's Hawthorne Works: illumination experiments; a “relay assembly test room” where six women worked under varying conditions of rest breaks and hours; interviews with thousands of workers; and observation of a “bank wiring room” of men.
Finding
As told in textbooks: productivity rose whatever was changed, even when lighting was reduced, because workers responded to being observed and paid attention to. This became the Hawthorne effect. The bank wiring observations found that workers enforced informal norms on output, sanctioning “rate-busters” who worked too hard.
Why it mattered
It launched the “human relations” school of management, which emphasised social and psychological factors at work rather than pure incentives.
Afterlife
When economists Steven Levitt and John List found and reanalysed the original illumination data in 2011, they found little evidence for the effect as described. Reanalyses of the relay room point to other explanations, including the replacement of two workers and changes in pay structure. The informal group norms finding, by contrast, has been repeatedly confirmed in workplace studies.

Doubtful — as a demonstration of the Hawthorne effect. Observer effects do exist, but this study is a weak basis for them. The finding about informal work norms is worth more than the famous one, and is directly relevant to how workers respond to algorithmic monitoring.

The unemployed of Marienthal

Marie Jahoda, Paul Lazarsfeld, Hans Zeisel · 1933 · Marienthal, Austria · community study

Question
What does mass unemployment do to a community? Does it radicalise people, as many on the left expected?
Method
When the textile mill that employed almost everyone in the village of Marienthal closed around 1930, a team from Vienna studied the village in 1931–32 using an inventive mix of methods: time-use sheets, school essays, meal records, library lending figures, and even measuring how fast people walked down the street. To be useful rather than extractive, the researchers ran a clothing collection, a medical clinic and courses.
Finding
Not revolt but resignation. The researchers called it “the weary community”. Men lost their sense of time: they walked slower, stopped wearing watches, could not account for their days. Membership of political clubs fell; library borrowing dropped even though it was free. Families could be classified as unbroken, resigned, in despair, or apathetic, and as income support ran down, families slid towards apathy.
Why it mattered
It showed that work provides far more than income. Jahoda later developed this into latent deprivation theory: jobs supply a time structure, social contacts, participation in a collective purpose, status and identity, and regular activity.
Afterlife
Repeatedly confirmed in research on unemployment and well-being. Lazarsfeld went on to found modern American survey research.

Holds up Arguably the single most relevant study for the AI age. If machines replace work, Marienthal suggests that a basic income alone will not replace what work provides.

The People's Choice

Paul Lazarsfeld, Bernard Berelson, Hazel Gaudet · 1944 · Erie County, Ohio · panel survey

Question
How do people make up their minds in an election campaign? How powerful are the mass media?
Method
A panel of around 600 voters interviewed repeatedly, roughly monthly, through the 1940 presidential campaign, plus control groups. Panel designs, following the same people over time, were then novel.
Finding
Very few people changed their vote. Social characteristics (religion, class, place) predicted voting well. Media mostly reinforced existing views. Where influence happened, it often passed through opinion leaders: people who followed the media and passed ideas on in conversation. This became the two-step flow model of communication.
Why it mattered
It overturned the “hypodermic needle” view that mass media inject opinions into a passive public, and introduced the idea of “limited effects”.
Afterlife
Later research found stronger effects through agenda setting (what people think about) and framing. The opinion-leader idea was influential in marketing — it is the conceptual ancestor of the “influencer”.

Qualified Limited direct persuasion remains a robust finding, and it is a useful corrective to panics about AI-generated propaganda. But “limited” is not “none”, and the media environment of 1940 is not ours.

Goffman, Asylums

Erving Goffman · 1961 · St Elizabeths Hospital, Washington, DC · participant observation

Question
What is life like for the inmates of a large psychiatric hospital, seen from their own point of view?
Method
About a year (1955–56) of fieldwork in a federal hospital with more than 7,000 patients, working nominally as an assistant to the athletic director so that he could mix with patients without being identified with staff.
Finding
Mental hospitals share features with prisons, monasteries, barracks and boarding schools: they are total institutions, where all of life happens in one place under one authority. On entry, inmates undergo a “mortification of the self”: stripped of possessions, routines and identity. Many symptoms staff saw as illness were rational adaptations to the institution. Patients developed an “underlife” of small resistances and secondary adjustments.
Why it mattered
Alongside other critiques, it helped build the case for deinstitutionalisation — the closing of large psychiatric hospitals from the 1960s — though that policy's results, including homelessness among people with serious mental illness, were mixed.

Holds up As an ethnography and a concept. “Secondary adjustments” is the right term for how gig workers game the apps that manage them.

Milgram's obedience studies

Stanley Milgram · experiments 1961–62, published 1963 · Yale University · laboratory experiment

Question
In the shadow of the Eichmann trial: would ordinary people inflict harm on an innocent person simply because an authority told them to?
Method
Participants, told the study was about learning, acted as “teacher” and were instructed to give increasing electric shocks to a “learner” (an actor) for each wrong answer. The shocks were fake. An experimenter in a grey lab coat urged them on with set phrases. Milgram ran many variations.
Finding
In the best-known version, 26 of 40 men (65%) went all the way to the maximum 450 volts, past the learner's screams and then his silence. Obedience fell sharply when the victim was in the same room, when the experimenter gave orders by phone, or when other “teachers” refused.
Afterlife
The study triggered a revolution in research ethics, contributing to modern ethics review boards. Gina Perry's archival work (Behind the Shock Machine, 2012) showed that obedience varied enormously across the variants, that many participants doubted the shocks were real, and that the experimenter often went off-script. Jerry Burger's 2009 partial replication, stopped at 150 volts for ethical reasons, found rates similar to Milgram's at that point. Alex Haslam and Stephen Reicher argue the participants were not blindly obedient but identified with the scientific project — “engaged followership”.

Qualified People do defer to authority to a disturbing degree, but the situation matters hugely and the famous 65% is one number among many. Directly relevant to how people defer to machine recommendations: see “automation bias” in Concepts.

The Coleman Report

James S. Coleman and colleagues · 1966 · United States · national survey

Question
Required by the Civil Rights Act of 1964: how unequal were educational opportunities for children of different races, and did school resources explain differences in achievement?
Method
Equality of Educational Opportunity surveyed more than half a million pupils and their teachers in several thousand schools — one of the largest social science studies ever conducted at the time.
Finding
Differences in school spending and facilities explained surprisingly little of the variation in achievement. Family background mattered most, and the social composition of the school — the backgrounds of a student's classmates — mattered more than resources.
Why it mattered
The finding about school composition was cited in support of desegregation and busing. The finding about spending was cited by those who argued money would not fix schools. It also helped launch the economics and sociology of education production.
Afterlife
Later research using better data and methods has found that school spending does matter, particularly for low-income children. Family background remains the strongest predictor.

Qualified A landmark whose most-quoted conclusion was too strong. A model case of how one study can be used by both sides of a policy fight.

The Stanford Prison Experiment

Philip Zimbardo · August 1971 · Stanford University · simulation

Claim
Twenty-four male students randomly assigned to be guards or prisoners in a mock prison in a university basement quickly took on their roles; guards became abusive and prisoners broke down. The planned two-week study was stopped after six days. The lesson, repeated in textbooks for decades: situations, not dispositions, make people behave cruelly.
Problems
Thibault Le Texier's archival investigation (book 2018; article in American Psychologist 2019) showed that guards were given instructions and encouragement to be tough; that Zimbardo acted as prison superintendent rather than neutral observer; that at least one prisoner's famous breakdown was, by his own later account, partly faked to get out; and that the study was never designed or published as a normal experiment. The BBC Prison Study (2002) by Reicher and Haslam, run under controlled conditions, found very different dynamics: guards were reluctant to impose authority and prisoners organised.

Doubtful It is better understood as a demonstration of what people do when a leader encourages them to behave badly — which is an important finding, but not the one that was taught.

The strength of weak ties

Mark Granovetter · 1973 (AJS); Getting a Job, 1974 · Newton, Massachusetts · survey, network theory

Question
How do people find jobs, and which social connections help most?
Method
Interviews and a survey of professional, technical and managerial workers who had recently changed jobs, combined with a theoretical argument about network structure.
Finding
Many people found jobs through personal contacts — and most of those contacts were people they saw only occasionally or rarely, not close friends. Close friends tend to know the same people and hear the same news. Acquaintances bridge to other social circles and bring new information. Hence the “strength” of weak ties.
Afterlife
One of the most cited papers in all of social science. In 2022, a study in Science by Karthik Rajkumar and colleagues analysed experiments LinkedIn had run on its “People You May Know” feature, affecting some 20 million users over five years. It found causal evidence that weak ties helped job mobility — with the most useful ties being moderately weak, not the weakest.

Holds up And the 2022 study raised a new question. A platform's algorithm had been changing millions of people's networks, and their job prospects, as an experiment they did not know they were in.

“On Being Sane in Insane Places”

David Rosenhan · 1973 (Science) · US psychiatric hospitals · covert field experiment

Claim
Eight sane “pseudopatients” got themselves admitted to twelve psychiatric hospitals by reporting hearing voices, then behaved normally. None was detected by staff; most were diagnosed with schizophrenia and stayed an average of 19 days. In a follow-up, a hospital warned to expect pseudopatients identified dozens of genuine patients as fakes, though none had been sent.
Impact
Enormous. It fed into the anti-psychiatry movement and the overhaul of psychiatric diagnosis in the DSM-III (1980).
Problems
The journalist Susannah Cahalan's investigation (The Great Pretender, 2019) found discrepancies between Rosenhan's private records and the published paper, including in his own symptoms on admission. She could identify only two of the other pseudopatients; one had a positive hospital experience and appears to have been excluded from the data.

Doubtful The concern it raised — that labels shape how every subsequent behaviour is interpreted — is real and well documented elsewhere. This paper is no longer good evidence for it.

Bourdieu, Distinction

Pierre Bourdieu · 1979 · France · survey (1963, 1967–68), correspondence analysis

Question
Is taste in art, music, food and furniture a matter of personal preference, or does it follow social position?
Method
A survey of more than a thousand people on their tastes and practices, analysed with correspondence analysis, a technique that maps relationships between categories in a two-dimensional “social space”, combined with interviews and photographs.
Finding
Taste tracks class closely, and along two dimensions: the total volume of capital people hold and its composition — whether it is more economic (business owners) or more cultural (teachers, artists). Taste is a weapon of distinction: the dominant classes define their own tastes as refined and others' as vulgar, and this helps reproduce inequality across generations through the school system.
Afterlife
American researchers, notably Richard Peterson in the 1990s, found that high-status people had become “cultural omnivores”, consuming both high and popular culture; snobbery became breadth rather than exclusivity. Others argue omnivorousness is itself a new form of distinction.

Qualified The specific map is of 1960s France. The insight endures, and recommendation algorithms are now one of the main forces shaping taste — while learning its class patterns from data.

Hochschild, The Managed Heart

Arlie Russell Hochschild · 1983 · US · interviews, observation of training

Question
What happens when feelings become part of the job?
Method
Study of flight attendants, including observation of Delta Air Lines training, contrasted with bill collectors, whose job was to produce the opposite emotions.
Finding
Service work requires emotional labour: managing one's own feelings to produce a required emotional state in customers. Flight attendants were trained to smile sincerely, to think of passengers as guests in their living room, and to reinterpret rude passengers as frightened children. This “deep acting” could estrange workers from their own feelings — a new kind of alienation.

Holds up Emotional labour is now a core concept in the sociology of work. It frames two AI questions: what happens to workers when chatbots take over the routine emotional labour of customer service, and what happens to users when a machine performs care with no feelings behind it.

Moving to Opportunity

US Department of Housing and Urban Development; later Jens Ludwig, Lawrence Katz, Raj Chetty and others · 1994–present · five US cities · randomised experiment

Question
Neighbourhoods and life outcomes are strongly correlated. But is that because neighbourhoods shape people, or because different kinds of people end up in different places?
Method
Around 4,600 low-income families in public housing in Baltimore, Boston, Chicago, Los Angeles and New York were randomly assigned by lottery: some received vouchers to move only to low-poverty neighbourhoods, some ordinary vouchers, some nothing. Randomisation lets researchers separate neighbourhood effects from selection.
Finding
Early results were disappointing: no gains in adults' earnings. But adults who moved showed improvements in mental health and reduced rates of extreme obesity and diabetes. Then, in 2016, Chetty, Hendren and Katz followed the children into adulthood and found that those who moved before age 13 earned about 31% more in their mid-twenties and were more likely to go to college. Those who moved as teenagers did not benefit, and perhaps were harmed by disruption.

Holds up One of the most important experiments in modern social science. It shows that place has causal effects, and that they accumulate with time spent there. It has shaped housing voucher policy.

The mark of a criminal record — and name audits

Devah Pager (2003); Marianne Bertrand and Sendhil Mullainathan (2004) · Milwaukee; Boston and Chicago · field audit experiments

Question
How much do race and a criminal record affect a person's chances of getting hired?
Method
Pager sent matched pairs of young men — trained to present themselves identically, with equivalent fake résumés — to apply in person for around 350 entry-level jobs in Milwaukee, varying race and whether they reported a prison record. Bertrand and Mullainathan sent about 5,000 fictitious résumés in response to job ads, randomly assigning names associated with white or Black Americans (“Emily” and “Greg” versus “Lakisha” and “Jamal”).
Finding
Pager: white applicants with a criminal record received callbacks at 17%, slightly more often than Black applicants without one (14%). Bertrand and Mullainathan: résumés with white-sounding names received about 50% more callbacks.
Afterlife
Audit studies have been replicated across many countries and groups. A very large 2022 study by Patrick Kline, Evan Rose and Christopher Walters, sending some 80,000 applications to large US employers, found a smaller but persistent racial gap, heavily concentrated in a minority of firms. Pager's work contributed to “ban the box” policies, which remove criminal-record questions from initial applications — though some research suggests such policies can increase statistical discrimination against young Black men when employers lose information.

Holds up The audit study is sociology's best tool for measuring discrimination — and the direct ancestor of how researchers now audit algorithms: hold everything constant, vary one attribute, and watch the output.

MusicLab

Matthew Salganik, Peter Sheridan Dodds, Duncan Watts · 2006 (Science) · online · web experiment

Question
Why are hits so unequal and so hard to predict? Is success determined by quality, or by social influence?
Method
About 14,000 participants were invited to a website where they could listen to and download 48 songs by unknown bands. They were randomly assigned either to an independent condition, where they saw no information about others' choices, or to one of eight separate “worlds” where they could see how many times each song had been downloaded by others in their world.
Finding
Social influence made success both more unequal (the top songs captured far more downloads) and more unpredictable (the same song could be a hit in one world and a flop in another). Quality mattered at the extremes — the best songs rarely did very badly and the worst rarely did very well — but almost any other result was possible.

Holds up A foundational result for the age of platforms. Every recommendation system that shows popularity counts is running MusicLab at global scale, and inheriting its arbitrariness.

Obesity spreads through social networks

Nicholas Christakis, James Fowler · 2007 (New England Journal of Medicine) · Framingham Heart Study · longitudinal network data

Claim
Using records on more than 12,000 people followed from 1971 to 2003, the authors reported that a person's chance of becoming obese rose by 57% if a friend became obese, and that effects extended to friends of friends and even friends of friends of friends (“three degrees of influence”). Similar studies followed on smoking, happiness and loneliness.
Dispute
Statisticians including Cosma Shalizi and Andrew Thomas showed that, in observational network data, contagion is generally impossible to separate from homophily (people befriend people like them) and shared environment. Other analyses found that similar “contagion” could be detected for traits like height, which cannot spread socially.

Contested Social influence on health behaviour is plausible and supported by experimental work; the specific claims about multi-step contagion from observational data are not established. A useful lesson in why correlations in network data, the raw material of social platforms, are so hard to interpret.

The 61-million-person experiment

Robert Bond, James Fowler and colleagues, with Facebook · 2012 (Nature) · US congressional elections, 2010 · platform experiment

Question
Can online social messages change real-world political behaviour?
Method
On election day 2010, about 61 million Facebook users were shown a message encouraging them to vote with an “I Voted” button. For most, it also showed pictures of friends who had clicked it. Smaller groups saw the message without friends' faces, or no message. Voting was checked against public voter records for a subset.
Finding
The informational message alone had essentially no effect on turnout. The social version did, both directly and through friends of those who saw it. The authors estimated about 340,000 additional votes in total, most of them through social contagion between close friends.

Holds up Small per-person effects become large at platform scale. The study also demonstrated a power that raised immediate concern: a single company could, in principle, nudge turnout selectively.

Facebook emotional contagion

Adam Kramer, Jamie Guillory, Jeffrey Hancock · 2014 (PNAS) · Facebook, January 2012 · platform experiment

Method
For one week, the News Feeds of 689,003 users were altered to show fewer posts with positive or with negative emotional words, and the emotional content of their own posts was measured.
Finding
Users who saw fewer positive posts wrote slightly fewer positive words and slightly more negative ones, and vice versa. The effect was statistically significant because the sample was huge, but extremely small in size.
Controversy
Users had not been asked. The company argued its data-use policy covered the research. Many researchers, and the public, disagreed. PNAS published an “Editorial Expression of Concern” noting that the study may not have met the principles of informed consent and opt-out.

Qualified The finding is tiny; the controversy was large and important. It made visible that platforms experiment on their users constantly, and it pushed the field to rethink research ethics for the digital age. Salganik's Bit by Bit treats it as a central case.

“Machine Bias” and the COMPAS debate

Julia Angwin, Jeff Larson, Surya Mattu, Lauren Kirchner (ProPublica) · May 2016 · Broward County, Florida · data journalism and audit

Question
Is a widely used commercial risk-assessment tool, which scores criminal defendants on their likelihood of reoffending, fair across races?
Method
ProPublica obtained COMPAS scores for more than 7,000 people arrested in one Florida county in 2013–14 and checked who was actually charged with new crimes over the next two years.
Finding
Black defendants who did not reoffend were almost twice as likely as comparable white defendants to have been labelled higher risk (about 45% vs 23%). White defendants who did reoffend were more often labelled low risk. The company replied that its scores were equally accurate for both groups: a given score meant about the same probability of reoffending whatever the defendant's race.
The twist
Both were right. Computer scientists, including Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan, and separately Alexandra Chouldechova, proved that when two groups have different underlying rates of the outcome, it is mathematically impossible for a score to satisfy both definitions of fairness at once. And in 2018, Julia Dressel and Hany Farid showed that untrained people recruited online predicted reoffending about as accurately as COMPAS, as did a simple model using just age and number of prior convictions.

Contested — and instructive The deepest lesson is sociological, not technical: “fairness” is not a property you can compute once and for all. Choosing between definitions is a political decision, and the different base rates the models inherit are themselves the product of policing, poverty and history.

Desmond, Evicted

Matthew Desmond · 2016 · Milwaukee · ethnography plus original survey

Question
What role does housing, and the loss of it, play in the lives of the urban poor?
Method
Desmond lived in a trailer park and then a rooming house in Milwaukee's inner city in 2008–09, following eight families and two landlords, and combined this with a survey of around 1,000 renters and court records.
Finding
Many poor renting families spent more than half, often far more, of their income on rent. Eviction was common and especially frequent for Black women, and for families with children. And eviction was not only a result of poverty but a cause of it: it led to job loss, worse housing, depression and school disruption.
Afterlife
Won the 2017 Pulitzer Prize for general nonfiction. Desmond founded the Eviction Lab at Princeton, which built the first national database of evictions in the US; its data became essential during the COVID-19 eviction moratoria. The book changed housing policy debates in several cities.

Holds up A model for combining ethnography with data. Also relevant to AI: tenant-screening algorithms that use eviction records can lock people out of housing for years over filings that never led to a judgment.

Gender Shades

Joy Buolamwini, Timnit Gebru · 2018 · MIT Media Lab · algorithm audit

Question
Do commercial facial analysis systems perform equally well across gender and skin type?
Method
The researchers built a balanced benchmark of 1,270 photos of parliamentarians from three African and three European countries, labelled by gender and by skin type on a dermatological scale, and tested gender classification products from IBM, Microsoft and Face++.
Finding
All performed worst on darker-skinned women, with error rates up to 34.7%, compared with at most 0.8% for lighter-skinned men. The gap only became visible when gender and skin type were analysed together, an intersectional analysis in the most literal sense.
Afterlife
A follow-up audit in 2019 found the named companies had substantially reduced their errors. In 2020, after the protests following George Floyd's murder, IBM announced it would leave the general-purpose facial recognition business, and Amazon and Microsoft paused sales of facial recognition to police.

Holds up It showed that public auditing can change corporate practice, and that what is not measured is not fixed.

Bias in a hospital algorithm

Ziad Obermeyer, Brian Powers, Christine Vogeli, Sendhil Mullainathan · 2019 (Science) · a large US health system · algorithm audit

Question
A widely used commercial algorithm identifies patients with complex health needs for extra care programmes. Is it fair to Black patients?
Method
The researchers obtained the algorithm's risk scores and the actual health records of tens of thousands of patients at one academic hospital, and compared how sick Black and white patients were at each score.
Finding
At any given risk score, Black patients were considerably sicker than white patients. The cause was not race in the data; the algorithm did not use race. It was the choice of target. The algorithm predicted future health costs as a proxy for health needs, and because less money had historically been spent on Black patients with the same conditions — through unequal access and treatment — it learned that they needed less care. Correcting the bias would have raised the share of Black patients automatically selected for extra help from 17.7% to 46.5%.
Afterlife
The authors worked with the manufacturer, which was able to reduce the bias substantially by changing what the model predicted. Similar algorithms were estimated to affect care decisions for some 200 million people in the US each year.

Holds up Perhaps the clearest single demonstration of how social inequality enters a model invisibly, through a seemingly neutral choice of what to measure. A sociologist would call it the problem of operationalisation.

The Fragile Families Challenge

Matthew Salganik, Ian Lundberg, Alexander Kindel, Sara McLanahan and 100+ co-authors · 2020 (PNAS) · US · mass prediction challenge

Question
With rich data and modern machine learning, how well can we predict how individual children's lives will turn out?
Method
The Fragile Families study had followed about 4,000 children born in large US cities around 2000, collecting thousands of variables from birth to age nine. Researchers held back data on six outcomes at age 15 (such as school grades, eviction of the household, and material hardship) and invited teams to predict them. 160 teams submitted predictions using everything from simple regressions to sophisticated machine learning.
Finding
Even the best predictions were not very accurate, and were only slightly better than a simple benchmark model using four variables. The teams' predictions were also similar to each other: the families that were hard to predict for one team were hard to predict for all of them.

Holds up A sobering result for anyone who believes enough data will predict individual lives. It is a strong argument for humility when algorithms are used to make decisions about specific people — children in social services, for instance.

Social capital and economic mobility

Raj Chetty, Matthew Jackson, Theresa Kuchler, Johannes Stroebel and colleagues · 2022 (Nature, two papers) · US · de-identified data on 72 million Facebook users

Question
Which kinds of social capital actually help people move up the income ladder?
Method
Using de-identified data on 21 billion friendships among 72 million users aged 25 to 44, the researchers measured three types of social capital for every US county, ZIP code, high school and college: connectedness between people of different incomes, network cohesion, and civic engagement.
Finding
Economic connectedness — the share of friends with high incomes among people with low incomes — was one of the strongest predictors of upward mobility yet found, stronger than school quality, job availability or family structure. The second paper showed that cross-class friendships depend both on exposure (whether rich and poor people are in the same places) and on “friending bias” (whether they actually befriend each other when they are).

Holds up (correlational) The associations are large and robust; causal claims rest on additional evidence. The study is also a case of a sociological question — Granovetter's and Putnam's — answered using a platform's private data, a reminder that much of the most important data about society now belongs to companies.

The feed and polarisation

Christopher Bail and colleagues (2018); Hunt Allcott and colleagues (2020); the US 2020 Facebook and Instagram Election Study (2023) · field and platform experiments

Question
Do social media and their algorithms cause political polarisation?
Bail, 2018
Twitter users in the US were paid to follow a bot that retweeted messages from the other side for a month. Rather than becoming more moderate, Republicans became significantly more conservative; Democrats became slightly more liberal, though not significantly. Exposure to the other side can backfire.
Allcott, 2020
Paying people to deactivate Facebook for four weeks before the 2018 US midterms reduced polarisation on policy issues somewhat, increased subjective well-being slightly, and reduced their knowledge of the news.
Meta, 2023
Academic researchers worked with Meta on experiments during the 2020 US election, published in Science and Nature in July 2023. Replacing the ranked feed with a chronological one, removing reshared content, or reducing exposure to like-minded sources by about a third changed what users saw substantially — but had no measurable effect on their political attitudes or polarisation over three months. Critics noted the short window, the unusual election period, and that Meta had altered its systems during the study; Meta's involvement raised questions of its own.

Contested The best evidence suggests algorithms shape what people see far more than they shift what people believe in the short term. Long-term, society-wide effects are much harder to test, and the strongest form of the “echo chamber” claim is not supported.

Generative AI at work

Erik Brynjolfsson, Danielle Li, Lindsey Raymond; Shakked Noy, Whitney Zhang; Fabrizio Dell'Acqua and colleagues · 2023–25 · field and lab experiments

Question
What does generative AI actually do to workers' productivity, and to whom?
Customer support
Brynjolfsson, Li and Raymond studied more than 5,000 customer-support agents at a software firm as an AI assistant was rolled out. Productivity (issues resolved per hour) rose by about 14% on average, by about a third for the least experienced and least skilled workers, and barely at all for the most experienced. The AI appeared to spread the tacit know-how of top workers to novices. (Published in the Quarterly Journal of Economics, 2025.)
Writing tasks
Noy and Zhang (Science, 2023) gave college-educated professionals writing tasks; those with access to ChatGPT finished about 40% faster with higher-rated quality, and the gap between weaker and stronger writers narrowed.
The jagged frontier
Dell'Acqua and colleagues (2023) worked with consultants at Boston Consulting Group. On tasks inside the AI's capabilities, consultants with AI were faster and produced higher-quality work. On a task just outside them, consultants with AI were less likely to get the right answer than those without, because they trusted plausible but wrong output.

Qualified (early) Consistent early finding: gains are largest for less experienced workers, which could compress skill differences within jobs. But the studies are short-term, and say little about wages, employment levels or who captures the gains — the questions sociology of work will be asking for the next decade. See The algorithmic society.


Reading across the studiesFive lessons

  1. Famous is not the same as sound. Hawthorne, Stanford and Rosenhan were among the best known studies in the social sciences, and are among the weakest. Popular summaries, and the AI systems trained on them, often lag decades behind the critical literature.
  2. The best designs vary one thing at a time. Moving to Opportunity, the audit studies, MusicLab and the algorithm audits all work because they isolate a cause.
  3. Scale changes the meaning of small effects. A tiny nudge to 61 million people is an election-sized effect.
  4. What you choose to measure is a moral decision. The hospital algorithm measured cost and called it need.
  5. Individual lives remain hard to predict. The Fragile Families Challenge should be required reading for anyone deploying a model to make decisions about people.

Questions for discussion

  1. Why might weak studies like the Stanford Prison Experiment become more famous than strong ones?
  2. The Facebook contagion and LinkedIn weak-ties studies both used experiments that platforms ran on their users. Should such data be used in science? Under what conditions?
  3. Apply the lesson of the hospital algorithm to another domain: what proxy might an AI hiring system or school admissions system use that hides a social inequality?
  4. What would a “Marienthal for the AI age” study look like? Design one.