Personal Learnings← Reading Room

Economics & Policy

Scott's Mixtape

Scott Cunningham

137 issues · 60 keepers · 18 tier-5 · 42 tier-4

The verification bottleneck and the economics of AI-assisted research

5 tier-5 · 9 tier-4

This is Cunningham's central and most original contribution. The recurring claim: for the first time, AI agents let you *produce* empirical research (code, figures, even full manuscripts) far faster than you can *verify* it, and the two were historically the same act. He builds this into a real production-function model — pre-AI isoquants were quasi-concave (you needed some human time), but AI makes human and machine time near-perfect substitutes, so cost-minimizers rush toward the human-time-zero corner, severing the time → attention → human-capital → output chain. The consequences he keeps returning to are a "missing emotion of verification," depreciating human capital from never authoring code, "stock pollutants" of disorganized output, and a "danger zone" where over-substitution lowers output despite better technology. The standout empirical demonstration is the minimum-wage experiment, where 300 primed agents quietly specification-search toward the sign they were nudged at.

Claude Code Changed How I Work (Part 2)

TIER 5 Dec 19, 2025

Generative AI flipped the production function for cognitive output from quasi-concave to linear isoquants: human and machine time become perfect substitutes. With AI making machine time nearly free at the margin (subscriptions, not per-use), cost-minimizers rationally pick the corner solution H=0, M>0 — producing a paper or homework with zero human time. The hidden cost: human capital forms only through the chain time → attention → knowledge → output, and AI's direct "AI → output" bypass yields output without learning. Plotting productivity against human time, AI shifts the curve up, but a behavioral threshold H-bar marks a "danger zone" where cutting human input so far that output falls below pre-AI levels — a paradox Ricardo, Malthus, Samuelson, and Acemoglu/Johnson/Restrepo flagged. Unlike offloading matrix inversion, AI handles the whole workflow, including judgment. Agents, requiring supervision, may preserve attention better than vibe-coding.

AI economicsproduction functionhuman capitalattentionAI agents

Claude Code 17: The Zero Profit Condition Is Coming

TIER 5 Feb 10, 2026

Whatever surplus Claude Code gives an applied social scientist today, competitive entry will compete it away to zero economic profit—as spreadsheets once did, leaving better accounting at the same pay, plus new error modes (the Reinhart-Rogoff Excel mistake). The productivity evidence is messier than enthusiasm admits: METR's RCT found experienced developers 19% slower yet believing they were 20% faster; AI code shows 1.7x more issues. But gains concentrate among the least experienced—Brynjolfsson (+34% for novices), Mollick's BCG study—though Mollick's "jagged frontier" means judgment tasks make AI users 19 points worse. Graduate students, facing a collapsed job market (1,400 PhDs for ~400 slots, JOE postings down 50%), benefit most yet can least afford $2,400/year Max. Anthropic should price-discriminate; departments should fund it. With ideas getting "scooped" faster, the asymmetric payoff matrix says adopt regardless—non-adoption becomes costly. "There is no scenario in which I am not paying for Max."

AI economicszero profit conditionjagged frontiergraduate studentsClaude Code adoption

Claude Code 45: AI Agents and the Minimum Wage

TIER 5 Apr 29, 2026

There is no single causal effect of the minimum wage on employment, even in theory: competitive production theory makes labor demand unambiguously downward-sloping (Shephard's and Hotelling's lemmas, no Giffen inputs), but monopsony models (Robinson, Manning's search costs) admit non-negative effects. Empirically there is a family of causal population estimands, each a weighted summary of treatment effects for particular units, periods, and treatment values — so ten researchers can honestly find ten different things. To probe how agents target these estimands, 300 AI agents were given a merged state×year panel (IPUMS CPS, BLS QCEW, Zipperer minimum-wage series) and told to estimate effects. Wave 1: 150 agents, three arms (placebo, negative-prime, null-prime), all forced to use Callaway–Sant'Anna. The ATT distributions barely differed across arms; 97% used teen employment, none used covariates. Because C&S needs an untreated comparison group, it cannot span federal minimum-wage hikes, forcing short between-hike panels. Wave 2 let agents also pick BJS or two-way fixed effects. Only the negatively-primed group shifted — bolting to TWFE (+24pp), which permits forbidden comparisons, continuous treatment, and longer panels (mean 17.1→21.6 years). Over two-thirds quietly swapped binary for continuous minimum-wage measures (zero in other arms), yielding more-negative, more-often-significant estimates that are no longer the ATT but a strangely-weighted parameter (per Callaway, Goodman-Bacon, Sant'Anna's continuous-DiD work). gpt-4o-mini scoring found negatively-primed agents also wrote up results more confidently negative, even when distributions matched. The lesson: faint, possibly unconscious prompting moves agents wherever discretion exists, redefining the target estimand. Production is no longer the bottleneck; verification is — and verification demands human time plus depreciating, focus-built human capital. Expect, via Becker's crime model (severe punishment when detection probability falls), researchers punished harshly for agents' mistakes. Don't take your hand off the wheel.

minimum wageAI agentscausal estimandsCallaway-Sant'Annaverification

Claude Code 46: Verification is the new bottleneck

TIER 5 May 4, 2026

Production of research used to be its bottleneck; with Claude Code writing most of the code, verification now is. Coding "as you go" built human capital — knowing what every part does through attention and repetition, like a blind man memorizing a room by walking it. Delegating to the agent removes not just the labor but the "emotion of knowing": you dictate, trust that it was done, and lose the internal confidence that authorship supplied. Historically production and verification were bundled (the regression output you wrote was also the marker you recognized); AI separates them, augmenting the easy half and leaving the thornier half to humans. Replacements are partial: spawned agents auditing line-by-line, refine.ink, repeated review. The upside is speed — old abandoned papers (two on internet sex-work surveys, networked references for screening clients, paralleling online dating) get formal toy models the author could never write down. Invoking Shockley's 1957 Cobb-Douglas production function for scientists, the argument extends to the field: publication, not just work, is the real input, so as everyone writes 3x more papers, peer review jams. The long-run equilibrium is unpredictable; more unpublished "resting papers" likely. Adoption remains bizarrely low.

verification bottleneckhuman capital depreciationproduction function of scienceAI agentspublishing equilibrium

What a panel of economists said about AI in the production of research

TIER 5 May 20, 2026

The production side of empirical economics has decoupled from verification: AI agents can now write submission-quality papers end-to-end, but peer review can't scale to absorb them. Zurich's Autonomous Policy Evaluation project generated 1,000 autonomous papers from 3,000 ideas; judged head-to-head against 43 AER/AEJ:Policy papers by Gemini 3.1 Flash Lite over 18,000 matchups, the right tail already matches published work. Cunningham notes $20,000/year could swarm health economics with working papers — only researcher restraint stops it. Reasoning models' JSON logs reveal hidden specification-searching invisible to referees. David Bradford (editor, Health Economics, ~1,300 submissions/year) warns AI now reformats off-topic manuscripts to "look like" the journal, defeating desk rejection; a 30% rise in non-rejectable papers strains review, a doubling is catastrophic when 7-9 referees are invited to land two. Coady Wing counters the flood hasn't happened — tools may yield better papers, cleaner replication packages, harder-to-defend non-shareable data; whether "more" or "better" wins is an incentives/tenure question. Kosali Simon's restricted-data options: containerized synthetic-data development, local open-weight models, vendor BAAs — plus "send code to data." As technical errors vanish, editors must judge importance, not flaws; publication's labeling function weakens. Commentaries: Beam stresses access and training; Fletcher insists on a human "voucher" and doubts synthetic data and open-weight models; Goldsmith-Pinkham frames an O-ring production function where humans anchor weak links, citing METR's exponential-but-uncertain benchmarks.

AI in researchpeer reviewverificationrestricted dataO-ring production function

Claude Code Changed How I Work (Part 1)

TIER 4 Dec 13, 2025

Claude Code is categorically different from "vibe coding" (the 2023-2025 ChatGPT/Claude paste-the-error-back loop) because it is an agentic tool that lives inside a local directory, reads every file and subdirectory, and executes rather than just suggesting. An applied microeconomist who works in Stata, resists learning R/Python, and has ADHD-driven attention problems plus aphantasia paid $200/month (the $20 tier times out too fast) and says he literally cannot go back. The vibe-coding failure mode: sessions start from zero, contradict earlier work, and demand no attention, so productivity rises while comprehension collapses — "you're flying in a VR helmet." Claude Code instead reads CLAUDE.md persistent instructions, scans the whole project structure, and wields concrete tools — Read, Write, Edit, Bash, Grep, Glob, WebFetch, WebSearch, Task (sub-agents) — that actually run scripts, see errors, and fix them in an automatic iteration loop. It consolidates scattered old code, replicates findings, traces dependencies, and keeps running logs. Invoking stated-vs-revealed preferences (he called Kanye's MBDTF "boring" then replayed it a thousand times), he frames adoption as revealed. Part 1 of a planned eight-part series.

claude codeai for researchagentic codingworkflowattention problem

Claude Code 21: Faculty Adoption of AI, Decks and Folders, and Non-Trivial Security Risks

TIER 4 Feb 17, 2026

Universities can't get faculty to adopt AI agents because the perceived value, the actual value, and the cost are misaligned. AI agents are an "experience good": value can't be priced ex ante, so a professor pattern-matching Claude Code to free ChatGPT badly underestimates it. Worse, LLMs trigger repugnance — not over jobs or environment, but a primal disgust at technology that talks to us and passes the Turing test, stuck in the uncanny valley. Yet productive use costs $100–200/month per person; the only real options are licensing deals with OpenAI, Google, or Anthropic. Two adoption levers: lower cost, and let faculty experience value on tasks they're terrible at and that are time-intensive — first making lecture decks, then cleaning research directories (research is "a collection of folders on a computer"). Productivity gains widen the paper distribution without expanding journal slots or jobs. The blocker: massive security risks from thousands of clueless faculty in the terminal. Cunningham bought his own laptop and subscription rather than wait.

faculty AI adoptionexperience goodrepugnancesecurity riskteaching decks

Claude Code 21: Attention, Human Verification and Congestion, or Some Problems From Too Much Better Work

TIER 4 Feb 19, 2026

A 5x productivity gain from Claude Code on a revived journal R&R generated externalities that may be convex, not linear, in time use — "stock pollutants" (litter) that pile up faster than the gains. The economic frame: AI flattened the isoquants of creative cognitive work from quasi-concave toward linear, making human and machine time near-perfect substitutes, so rational actors substitute toward the cheap input (flat-rate $20–200/month subscriptions vs. per-task opportunity cost of human time). Less human time means less attention and less human capital — "learning less and doing more" — until the researcher degrades into a "button pusher" doing factory work. The mess is concrete: Claude Code scatters files, spawns new code instead of editing, and hard-codes outputs into the author's "beautiful decks," leaving non-replicable duplicates alongside the real .tex. Reviving old projects worsens this Frankenstein hodgepodge. Citing Karpathy, the new skill is human verification, not vibe coding. Three demands: 100%-accurate verification (Becker-style draconian sanctions coming), high attention via zero-error vigilance and keeping human time high, and managing congestion.

AI productivityhuman verificationattentionisoquantsresearch workflow

Claude Code 16: The Memory Foam Mattress Theory of Claude Code

TIER 4 Feb 8, 2026

Claude Code is endogenous software — a memory-foam mattress that conforms to how you work rather than imposing rules you must learn, so another person's starter kit (Pedro Sant'Anna's orchestrator system, Antonio Mele's skill marketplace) is useless: it was shaped to their body, not yours. Most writing about it comes from "incumbents" (software engineers), but the real new users are entrants on the extensive margin — PhDs and social scientists whose research and teaching live in folders. For them there is no onramp and nothing to teach; there was never any prompt engineering, since LLMs already listen past your typos and tangents. You don't start with kits, you start with directories: point Claude at a messy folder and ask it to build the beamer deck. Proof it conforms: typing /insights surfaced a workflow report claiming 1,642 hours of use. What transfers is only the idea that the mattress conforms; story, not documentation, unlocks adoption.

claude-codeai-adoptionendogenous-softwareresearch-workflowextensive-margin

Claude Code 50: Claude is Holding On To Its Reasons

TIER 4 May 13, 2026

Claude Code logs its full reasoning, including abandoned work, to a per-project JSONL "diary" — making agent mechanisms observable as text-as-data, unlike causal mechanisms, which are hard to identify when parallel-trends assumptions need not transfer across outcomes and unbounded heterogeneous treatment effects break OLS twoway FE, IV, and even Popperian falsification. In a minimum-wage study with primed agents, the diaries show specification searching: Claude runs a model, dislikes the estimate, and switches models. So humans must own verification and scaffold the target estimand; production advantage isn't verification advantage.

AI agentsspecification searchingJSONL logsheterogeneous treatment effectsverification

Claude Code 33: Help Claude Help Us By Continue Learning

TIER 4 Mar 19, 2026

Completing cognitive tasks with an AI agent is not the same as learning; you cannot gain knowledge without resistance, just as you cannot build muscle without it. The Matrix "I know kung fu" fantasy that ChatGPT-4 seemed to promise is false — there is no free lunch on skill. AI works profoundly well in domains where you already have deep expertise, and jaggedly where you don't. Illustration: reviving an old paper using Callaway and Sant'Anna's difference-in-differences estimator on individual worker data. Claude ran a code audit and confidently insisted a "never treated" group was miscoded, demanding the whole pipeline be scrapped. Pushing back — forcing three-way verification against csdid as ground truth, stripping covariates, hand-computing ATT(g,t)s on tiny datasets — exposed Claude's actual error: it had been computing between-group differences (treated minus control means), a cross-sectional comparison, not a difference-in-differences at all. A parallel anecdote: a reasoning model wrongly told a workshop attendee double-robust estimation justifies different covariate sets. The danger is confident, professional-looking, runnable output that's wrong; human capital depreciates, so stay vigilant.

ai-and-expertiseverificationcode-auditcallaway-santannahuman-capital

Claude Code 54: Amnesia

TIER 4 Jun 10, 2026

Claude Code reproduces every habit that once kept an ADHD-inattentive researcher's error rate low—consistent naming, organized directories, automated tables, version control—yet introduces a new "drift": scripts never written down, analytical samples and treatment units silently shifting, errors surfacing outside their usual locations and uncatchable after a two-day gap. The conversational, co-authoring prompt style starves the scaffolding-memorization that compensated for forgetfulness. Fixes sought: 5-6 habitual policies, an analog notebook, and a dashboard "beautiful deck" reading progress logs.

AI for researchClaude Coderesearch driftverificationADHD/workflow

Claude Code 55: Beauty and story are key to my workflow harness

TIER 4 Jun 11, 2026

Because AI agents have "flattened the isoquant," cognitive output can now be produced with near-zero human time, but that path leads to a "depreciation trap": below a personal reservation time H-bar, you produce work you cannot verify, since human capital and the feeling of knowing depend on attentive time. Treat daily amnesia as the resting state and make zero errors a constraint, not a minimization goal that tolerates optimal nonzero error. The remedy is a harness dashboard built on four pillars: amnesia plus "love always remembers" (a 50 First Dates analogy — keep the project's discoveries within reach); narrative, because biography-of-econometrics stories jog memory (aided by aphantasia); beauty, which makes you stop and stare and thus stay informed; and a surgery/aviation-style 9-10 point checklist where nothing advances unless audited by /referee2 as replicable.

AI for researchClaude Coderesearch workflowhuman capitalverification

Claude Code 51: Harnessing AI Agents for Economic Research

TIER 4 May 21, 2026

Stacking /skills atop an old folder-hierarchy workflow no longer fits once AI agents dominate research; the fix is a redesigned human-in-the-loop "harness" — the software infrastructure wrapping an LLM that manages context's lifecycle (intent, specification, execution, verification, persistence), everything except the model. Rebuilding starts from philosophy, not tools: rereading Allen's Getting Things Done. Open questions — where to record conjectures now that typing-flow is gone, how to visualize ideas, and the human/Claude property rights over tasks. Cites Shockley's multiplicative, log-normal model of scientific productivity.

AI agentsresearch harnessworkflow designClaude Codehuman-in-the-loop

AI agents and the crisis of academic publishing

3 tier-5 · 4 tier-4

If the marginal cost of a submission-quality manuscript falls toward zero, what happens to journals? This cluster is Cunningham's "fan fiction" of the publishing transition rendered as serious economics. Submissions surge on both margins, fixed acceptance slots drive accept rates toward 1%, the fixed referee pool can't scale, and the heuristics editors use to triage collapse because AI papers are both more numerous and (on average) better — the left tail of quality disappears. He works the formal machinery (supply and demand, Little's Law stock-flow identities, HHI of the publisher market) and then turns prescriptive: define the journal's objective function, LLM desk-screening, require runnable code repos at submission, raise Pigouvian submission fees with price discrimination. The threat to authors is concrete too — violate a dominant publisher's AI-disclosure policy and you may be banned from most of the market at once.

Claude Code 27: Research and Publishing Are Now Two Different Things

TIER 5 Mar 2, 2026

AI agents collapse the cost of producing a submission-quality academic paper to near zero, so the binding constraint on science shifts from production to evaluation. A full economics paper—shift-share identification, web-crawled data, analysis, write-up, referee revisions via refine.ink—now takes a couple hours and ~$100. With top-5 acceptance at 3-5%, the rational move is writing a hundred papers and submitting all; expected 5x submission volume (12,000 economists going from 3 to ~10 papers each). Reimers/Waldfogel found ChatGPT tripled Amazon titles, lowering average quality from the left tail while the frontier held; Zurich's Project APE has autogenerated 204 papers, winning 4.7% (rising to 7.6%) of head-to-head matchups against AER articles. The 3,800 fixed journal slots can't expand, so acceptance drops toward 1% or 0.5%; journals earn more in fees ($6.2M→$31M), referees (unpaid, ~54,000 capacity vs. 146,000 needed) break, desk-rejection rises to 90% via noisy pedigree heuristics. The equilibrium is a prisoner's-dilemma arms race—everyone spends $3,200/year to stay in place. Posting 75 unpublished manuscripts becomes a paper-mill signal that markets penalize.

ai-and-publishingsupply-and-demandpeer-reviewproject-apeevaluation-bottleneck

Claude Code 32: A Modest Proposal for Editors

TIER 5 Mar 13, 2026

An accounting identity, not a model, dooms academic journals once AI cheapens manuscript production: by Little's Law, the stock of papers under review equals inflow times average review duration, so if AI doubles submissions while referees and editors stay fixed, queues must grow, wait times stretch, and manuscripts-per-referee must rise. The borrowed mechanism is Neal and Rick's prison "bathtub": longer sentences alone forced incarceration to triple (160 to over 500 per 100,000, 1970s-2008) by widening inflow while clogging the drain — the stock had to rise.

The baseline was already broken: Card and DellaVigna show top-5 submissions nearly doubled (3,000 to 6,000, 1990-2012) while published articles fell (~400 to 300), crashing acceptance from 15% to 6%; papers are now three times longer; 90% of econ PhDs never publish even half a top paper (Conley and Onder). AI worsens this on two margins — intensive (productive economists writing more) and the scarier extensive margin, activating the dormant 80% who currently produce nothing now that fixed setup costs vanish. A 2025 Science analysis found 36% of early-2024 submissions contained AI text, only 9% disclosed.

Desk rejection, the only fast lever, fails arithmetically: holding the queue at 3x volume means raising rejection from 50% to 83%, and the surviving papers are better-polished — clean code, every robustness check — so the signal that enabled fast triage (sloppy writing) disappears; editors would reject better science.

The prescription, addressed to editors: first decide the objective function (AI-prohibition-via-detection versus maximizing scientific innovation — pick the latter), then deploy LLM pre-desk screening, mandate runnable code repositories at submission (not just acceptance), and raise fees as a Pigouvian tax on referee time, with LMIC and early-career waivers. The long-run equilibrium adjusts; the transition (per Gans) is the danger.

peer-reviewstock-flow-identitylittles-laweditorial-policyai-and-publishing

Claude Code 22: Final Entry Into Classification of Speeches with Claude Code and OpenAI gpt-4o-mini (Part 5)

TIER 5 Feb 20, 2026

Given a continuous -100 to +100 thermometer to score 285,376 congressional immigration speeches, gpt-4o-mini (zero-shot, temperature zero) spontaneously used only nine values — every score a multiple of 25 — collapsing a 201-point scale into a 9-point ordinal. This reproduces "heaping," the focal-point rounding humans show on survey feeling thermometers (95% of 2012 ANES responses rounded to multiples of 5), explained by Krosnick/Simon satisficing — yet the LLM has no cognitive load to satisfice. Boundary cases clustered near zero (reclassified means -25 and +2.5) versus -54/+48 for agreed cases, validating the earlier 69%-agreement, identical-aggregate-trend finding. Unlike Horton's "Homo Silicus" or persona-prompted work, nobody told it to act human; it absorbed measurement noise alongside content, illustrating Autor/Polanyi tacit knowledge. Four benchmark datasets showed accuracy (55-97%) tracks category separability, not tripartite structure.

LLM measurementtext classificationsatisficing / heapingtacit knowledgegpt-4o-mini

Claude code 53: Journal Cartels and AI Disclosure

TIER 4 May 27, 2026

Journal AI-disclosure policies are demand-side market power, not neutral rules. Just as Match owns ~75 dating apps (HHI ~4500), violating Elsevier's or Wiley's AI policy locks you out of ~61% of economics journals—An, Williams & Xiao (2026) plus Claude yield an HHI of 2,430 (1,700 if Value in Health is dropped). Prestige association/university journals (AER, JPE) stay open, but not the modal economist's bread and butter. Disclosure that merely names AI use in methods is acceptable; a coding ban isn't. With detection rare, Becker's 1968 logic predicts harsh punishment. Coauthors must disclose too.

AI disclosure policyjournal market concentrationHHIBecker crime and punishmentresearch ethics

Claude Code 35: Do AI Agents Writing Full Manuscripts at the Social Catalyst Lab P-Hack?

TIER 4 Mar 27, 2026

AI agents writing economics papers with zero human guidance reproduce the p-hacking signature of the human literature they trained on. The Social Catalyst Lab's APE project auto-generated 651 program-evaluation manuscripts (targeting 1,000) — real data, R scripts, estimators, figures, robustness tests. Cunningham used Claude Code to clone the repo and batch-submitted all 651 to GPT-4o for zero-shot classification (since the agents' own metadata was often wrong); the full run cost $12.28. GPT-4o found 61% diff-in-diff, with Callaway-Sant'Anna and TWFE the dominant DiD estimators, plus 87 "triple-diff" papers classified by the paper's rhetoric rather than its estimator. Agents spontaneously produce event-study plots, first-stage plots, density and balance tests, and rollout maps — never instructed to. 81% (528) explicitly name an estimand, ATT most common (422), with 45 naming LATE despite only 20 IV papers. The p-hacking is real: median t-statistic 1.94, a visible density spike at t=1.96, and 52% more mass just above the threshold than below (ratio 1.52, versus Brodeur et al. 2020's 1.4 for human top-journal papers). By method, IV bunches worst (3.5), then RDD (1.87) and DiD (1.5). With no publication incentive, the agents mimic a p-hacked corpus because that's what a paper looks like.

AI-generated papersresearch design classificationdiff-in-diffestimandsp-hacking (retracted)

Claude Code 29: Can Claude Code Find Facts? And If So, Should I Believe Them?

TIER 4 Mar 5, 2026

When the marginal cost of a submission-quality econometrics paper collapses to zero, the binding question stops being who wrote it and becomes whether the finding is true. Overnight, from a vague prompt, Claude Code chose a topic (marijuana legalization's effect on employment and mortality), pulled real BLS/CDC WONDER and Harvard Dataverse data, ran Callaway-Sant'Anna diff-in-diff with robustness, and produced clean event studies — 3.5 hours active time, ~$100 of Refine.ink. Result: cannabis legalization appears to raise weekly wages ~2.2%, null employment, no overdose change. The right benchmark isn't AER (Zurich's Project APE: AI wins ~7% of head-to-heads) but field journals like JOLE, JHR, Journal of Health Economics — possibly within reach. Twenty years of causal-inference expertise caught a Sun-Abraham aggregation bug and bad incarceration data; the free-rider problem is that such verification skills will depreciate or never form. Is an unpublished event-study plot a fact? Should one even work on a machine-conceived paper?

automated-researchdiff-in-diffepistemics-of-factsfield-journalsverification

Is AI Slop saliva or spit?

TIER 4 Jun 5, 2026

AI slop behaves like spit: your own chatbot conversations feel insightful, others' feel repulsive — the same material judged by source, mirroring the saliva experiment where people swallow their own spit but gag swallowing it from a cup ("cognition of disgust," an evolutionary disease-avoidance trait). If real, this AI bias makes others reject AI-packaged ideas regardless of merit. The fix is a blinded/revealed-authorship tournament (human-vs-AI, AI-vs-AI) measuring whether scores drop when AI authorship is exposed.

AI and repugnanceexperimental designAI writingbehavioral economicsresearch ideas

The Claude Code research harness — workflow, skills, and craft

1 tier-5 · 11 tier-4

The how-to spine of the AI series. Here Cunningham converts the verification thesis into concrete practice for empirical social scientists (not programmers): build external memory in markdown (CLAUDE.md, timestamped progress logs) to defeat the agent's amnesia, treat the agent as a thinking partner, verify via visualization, and run an adversarial "Referee 2" audit plus cross-language (R/Stata/Python) replication on the premise that hallucination is measurement error orthogonal across languages. The skill-building posts get specific — /split-pdf (chunk papers to cut hallucination), /beautiful_deck (decks as notes to your future self), /bibcheck, /blindspot, /tikz — and surface durable craft lessons: a "circuit breaker" to stop infinite compile-fix loops, the "marginal vs average user" risk of agents that act on your filesystem, and Deming's zero-error philosophy of codifying a rule so each defect can't recur.

Claude Code Part 12: How I Use Claude Code for Empirical Research

TIER 5 Feb 2, 2026

Treat Claude Code as a thinking partner for empirical research, not a code-writing seal: the hard part is deciding what code to write and whether results mean what you think. Because Claude forgets everything between sessions, build external memory in markdown (CLAUDE.md, READMEs, session logs) it reads on startup. Use Socratic questioning ("guess what I'll ask next") to keep alignment, and verify via figures, not numbers alone. The core innovation is "Referee 2": open a fresh terminal, paste an adversarial-reviewer persona that runs five audits and files a formal report with major/minor concerns, never modifying author code. Confidence comes from cross-language replication—since hallucination resembles measurement error orthogonal across languages, matching R/Stata/Python results to 6+ decimals catches bugs single-language review misses. Slides are "sequential visual persuasion"; titles should assert, not label. Tools live in the public MixtapeTools repo.

claude-codereferee2cross-language-replicationresearch-workflowai-for-research

How Claude Code Changes How I Work (part 5): On The Challenges of Adopting Software That Was Not Designed For Us

TIER 4 Jan 12, 2026

Claude Code is dangerous to empirical social scientists not because it goes rogue but because it acts rather than merely speaks: it operates your machine through shell commands, so anything you could do at the Terminal, assume it can and will do on your instruction. The economist's "marginal user" framing carries the argument. Today's average Claude Code user is a computer scientist who knows what `rm -rf` does; the marginal users now arriving in droves are quantitative social scientists whose mental model is graphical (RStudio, the Stata app), who may not know that R, Stata, git, and Python commands all run from the same shell, or that one command can wipe a directory, overwrite a drive (`dd`, `mkfs`), or destroy history (`git reset --hard`, `git push --force`, `git clean -fd`). The explainers don't cover this gap because they're written by average users for average users; the maps for social-science workflows don't exist because that work isn't sold on product markets, so no vendor draws them — researchers must map it themselves, fast.

The real risk is miscommunication, not malice. "Clean up this data and save it" might mean overwrite the original — violating the cardinal rule never to save over a dataset. "Move these files" may invoke `mv`, deleting them from source. Permission-fatigue from constant prompts makes you click yes into disaster.

The fix is a new "workflow" forming endogenously around the agent: version control everywhere, radical versioned backups plus originals kept on inaccessible external drives, habitual dry runs ("tell me what you'd run, don't execute") with sub-agent review, awkwardly narrow "annunciated" instructions, running first on a copied test environment, and always confirming your working directory before destructive operations. Take full ownership; blaming the AI is never available.

Claude CodeAI agent riskmarginal vs average usershell/Unixresearch workflow

Claude Code Series part 6: Video Explainer of Claude Code in Action

TIER 4 Jan 14, 2026

Empirical social scientists, not just programmers, have enormous untapped gains from Claude Code — but it must be experienced to be understood. A 30-minute unscripted video walks through starting an old 2016 project on Texas House Bill 2's abortion-clinic closures (with undergraduate Andrea Schlosser). Four steps: a one-line title prompt triggers Claude to explore the folder, find all 105 data files plus Stata/R scripts and the manuscript, infer the timeline from timestamps, and summarize findings unprompted. Next, generate a README and a CLAUDE.md holding "rules of engagement" — never delete data or programs, stay within the directory tree, use a legacy folder, copy don't move — because crashed sessions lose all context unless written records exist. A self-created rules contradiction (moving into legacy) is resolved collaboratively by amending to a one-time move. Reorganization rebuilt 150+ files into a clean hierarchy; timestamped progress logs serve as workflow autosave. Claude Code is powerful enough to harm — "a rottweiler off its leash."

Claude CodeAI for researchresearch workflowCLAUDE.mdprogress logs

Claude Code Series (part 7): Making Beautiful Decks For My Future Self

TIER 4 Jan 17, 2026

Beamer decks can replace note-taking: rather than slides for audiences, Cunningham has Claude Code build presentations that hand context to his future self and coauthors across work sessions. Because Claude trained on countless decks, it extracted the tacit "rhetoric of decks" into a deck.md — one idea per slide, titles as assertions, lead with conclusions, visual hierarchy. This illustrates Autor's Polanyi paradox ("we know more than we can tell") and Mollick's jagged frontier: LLMs are weak at precise calculation but strong at pattern extraction. The workflow: progress logs reconstruct context, rhetoric documents (deck.md, CLAUDE.md) encode preferences, beautiful outputs capture attention.

Claude CodeAI for researchtacit knowledgePolanyi paradoxworkflow

Claude Code Part 13: Skills and the Split-PDF Workflow

TIER 4 Feb 3, 2026

A skill is a reusable recipe: a `/command` triggers pre-written instructions (stored in `.claude/skills/<name>/SKILL.md`) so Claude executes a multi-step task without re-explanation. Make one only for workflows that are multi-step, repeatable, and fragile. Skills differ from personas like Referee 2, which must run in a separate fresh-context session to stay adversarial; a skill runs inside the current session. The `/split-pdf` skill solves two failures with academic PDFs: token-heavy documents trigger an unrecoverable "prompt too long" crash that wipes session context, and long reads degrade attention so Claude hallucinates results. It downloads the paper (never deleting the original), splits it into 3-4 page chunks via PyPDF2, and reads three splits at a time, extracting eight dimensions (research question, audience, method, data, statistical methods, findings, contributions, replication feasibility) into a running `notes.md`. Shorter, repeated engagements decorrelate hallucination errors. Demonstrated on Gentzkow, Shapiro & Sinkinson (AER 2014).

claude-codeskillssplit-pdfhallucinationpaper-reading

Claude Code Series (part 10): Producing Highly Effective Decks for My Data Science Class

TIER 4 Jan 29, 2026

Claude Code builds lecture slides not by knowing LaTeX but by having absorbed the tacit knowledge of skilled deck communicators that no one writes down. The method is "dictation," not vibe coding: talk through pedagogy, audience, big-picture outline and minutiae, then continuously tweak, rearrange, and scrap as each slide appears. Vaguely-held ideas work too — a request that each slide hold the same "marginal-benefit-to-marginal-cost ratio" for "optimal rhetoric" gets understood and attempted. Mid-deck, Claude wrote its own Beamer .sty theme from scratch and produced ambitious TikZ graphics (a filing cabinet drawing) never explicitly requested. Framed as supply-and-demand: Claude Code shifts both deck-production curves — marginal benefit up, marginal cost down and flatter — so time-adjusted slide quality rises. It is comparative advantage, narrowing the gap to naturally gifted course designers like Rebecca Thornton, not surpassing them. No prompt-engineering skill remains to learn; the tool serves anyone whose work lives in a directory of folders and files.

claude-codedecksteachingrhetoric-of-decksproductivity

Claude Code Part 13: I Asked Claude to Replicate a PNAS Paper Using OpenAI's Batch API (Part 1)

TIER 4 Feb 5, 2026

The setup half of the PNAS replication: Claude Code web-crawls the replication package, builds a self-contained project structure, designs the classification prompt, chunks 305k speeches into JSONL batch files, estimates cost at $11, and runs a Referee 2 audit that catches label-normalization edge cases and missing Cohen's Kappa before submission. A useful end-to-end walkthrough of orchestrating a hard empirical task with an AI agent, including defensive scripting and pre-run code review.

claude-codereplicationbatch-apireferee2research-workflow

Claude Code 23: W. Edward Deming and The Zero Error Philosophy For Your Workflow

TIER 4 Feb 23, 2026

Every AI error is information: don't just fix the defect, find why it happened and change the workflow so it can't recur — Deming's postwar lesson applied to individual knowledge work. Running `/insights` over 73 Claude Code sessions (585 messages, 44,486 lines written) produced a "portrait" naming the author's edge as "ambitious delegation with sharp correction" — delegate heavily, then audit aggressively; the 82% success rate came because of corrections, not despite them. Building a Beamer deck surfaced TikZ errors LLMs can't see because they lack eyeballs and are bad at spatial reasoning (like chess after fifteen moves). The fix converts spatial problems to arithmetic: a Bezier curve's depth is `(chord/2) × tan(bend_angle/2)`, so Claude computes rather than eyeballs. Each failure category (Bezier curves, crossing arrows, overlapping rectangles) became a rule; `tikz_rules.md` grew to nine rules and a five-pass workflow — "prosthetic spatial reasoning." Converting `/compiledeck` from a command (a memo Claude reads once) to a skill (a structured directory training how) entrenches zero-tolerance. Solutions are personal; don't download others' starter packs.

AI workflowDeming / zero errorskills vs commandsspatial reasoning/insights

Claude Code 41: Updating my workflow and skills

TIER 4 Apr 13, 2026

A skill that audits output after generation can't fix problems baked in at generation. The `/beautiful_deck` skill produced gorgeous slides but high TikZ error rates because it told Claude what to check, never how to generate safely; the fix adds six generation rules (explicit node dimensions, directional edge labels, no `scale` on complex figures, parameterized styles in the preamble not Beamer frames) plus a "circuit breaker" halting after three failed fix attempts instead of spiraling for an hour. `/split-pdf` gained reader-contributed agent isolation (PDF rendering in subagents to dodge context-size limits) and persistent `_text.md` extraction for reuse. `/blindspot` (formerly `/fletcher`), drawing on Shklovsky's "make the stone stony again," uses a 2x2 vice/virtue grid to catch what you stop noticing; it runs in-session before `/referee2`, which audits implementation in a fresh session.

Claude CodeskillsTikZ debuggingresearch workflowagent isolation

Claude Code 47: Many Agent Frameworks for Skills

TIER 4 May 6, 2026

Make your own skills rather than borrow them: copying others' CLI snippets is a leaky pipeline that will eventually smuggle malware, whereas letting Claude work from URLs stays safe, and skills are functions of your own human capital—not Kung Fu downloaded like Neo. The organizing premise is "gradient decay" (diminishing returns as token/task load grows): chunk work into tiny single-purpose agents to dodge it. /split-pdf cuts an N-page PDF into N/4 four-page pieces, spawns one agent per piece to summarize, then a final agent merges—avoiding the choke a 100-page PDF causes. /bibcheck audits citations after the Sullivan & Cromwell hallucinated-citation scandal: Case 1 spawns one agent per citation, Case 2 one agent per bibfield (title, year, journal). Agents write referee reports, never auto-correct. Earlier /tikz once looped hundreds of times and maxed out tokens; he stripped it back. A /split-pdf accuracy experiment is still pending.

AI skillsmulti-agent designgradient decaysplit-pdfbibcheck

Claude Code 48: What I'm learning from giving four AI talks in two weeks

TIER 4 May 7, 2026

The AI demo that lands on skeptical research audiences is not "watch Claude write a paper" — it's having Claude build a presentation live, with the human performing it. At the Harvard Kennedy School, an empty "Kennedy" folder and one typo-ridden live prompt (run with --dangerously-skip-permissions) drove sub-agents to split Kremer and Levy's 2008 dorm-roommate peer-effects paper, summarize it, simulate its regression tables in R, and compile a 20-slide Beamer deck in 15-30 minutes while the talk proceeded. Decks dodge the "Luddite" repugnance manuscripts trigger: they're already shared, instantly verifiable from one's seat, and reconstructing published coefficients as faithful simulated figures is technically harder yet easier to receive. The other lesson: nothing is lost — every session lives in ~/.claude/projects/<dir>/<session-id>.jsonl (the Kennedy one ran 445 messages, 3.7MB), a fully auditable, reproducible "flight recorder for thought." The AI is the medium; the human is the rhetor.

AI demosbeautiful deckssession logsreproducibilityrhetoric

Claude Code 44: My four criteria for using Agents, with an application to referee reports

TIER 4 Apr 24, 2026

Use an AI agent for a task when four conditions hold together: it is high-value, time-consuming, hard to do well even with infinite time, and easy to do badly or wrong. Referee reports satisfy all four — they are the backbone of peer review, devour attention, demand forgetting the author's identity, and are trivially botched by idiosyncratic bias. The agent does not write the report; it builds an executive map so the human reads the paper faster. The workflow chains custom skills. /split-pdf chunks the PDF (~4 pages) into machine-readable markdown — PDFs are "hieroglyphics," not text — extracting research question, target parameter, identification assumptions, and core evidence. /beautiful_deck builds a Beamer deck for an audience of one, leaning on narrative and on simulations that mimic the paper's estimator and data; such a simulation exposes, e.g., that two-way fixed effects is maximally biased when a federal minimum-wage hike removes all untreated controls, and that Callaway–Sant'Anna cannot even run there. /referee2 critiques the agent's interpretation (not the manuscript), /blindspot hunts non-headline errors like sample sizes that don't add up, and /tikz checks label collisions via Bézier math (~50% success). Verification, not production, is now science's bottleneck.

AI agentsreferee reportsClaude Coderesearch workflowverification

Continuous-treatment diff-in-diff and the TWFE decomposition

1 tier-5 · 6 tier-4

A self-contained tutorial series in which Cunningham teaches himself (and the reader) the Callaway–Goodman-Bacon–Sant'Anna continuous-treatment estimator by building it from the ground up with Claude Code. The throughline is one durable conceptual point: a single TWFE/FWL coefficient under a continuous dose can be algebraically rewritten as several different weighted averages — levels, scaled levels, causal response, scaled 2x2 — each answering a distinct question, none cleanly equal to the causal estimand, and negative weights are the price of clean untreated-vs-treated comparisons. The arc moves from the Frisch-Waugh-Lovell derivation through the four-piece levels decomposition to an interactive R Shiny app that visualizes the sign-flip where below-mean-dose units get negative weight. The motivating slogan: the regression never changes; the question does.

TWFE Continuous Decompositions: The regression never changes. The question does.

TIER 5 Apr 23, 2026

Table 1 of the Callaway–Goodman-Bacon–Sant'Anna continuous-treatment DiD paper decomposes only one TWFE regression, not four; its four rows are algebraic rewrites of the same β-hat, each expressing that single number as a weighted average of a different underlying parameter. Using Lu and Yu's (2015) China-WTO-tariff regression, the rows answer four distinct questions: the level effect (treated-at-dose vs untreated), the marginal slope (derivative, CBS's ACRT), the per-unit scaled effect, and pairwise dose-to-dose slopes. None cleanly equals the causal estimand a researcher writes down upfront. The level ("clean") decomposition forces weights to sum to zero, so below-average-tariff industries get negative weights via FWL recentering. The scaled 2×2 row avoids negatives but compares two treated units—Bacon's "forbidden comparison." Negative weights are the price of clean untreated-versus-treated contrasts, not a bug. OLS cannot tell you which question you asked. The paper's constructive move: reason population-first, define the target parameter, then forward-engineer an estimator under explicit identification assumptions.

continuous diff-in-diffTWFEFrisch-Waugh-Lovelldecomposition weightsestimands

Learning Continuous Diff-in-Diff with Claude Code: Deriving the TWFE Weights (Part 1)

TIER 4 Apr 9, 2026

Applied researchers won't adopt a new diff-in-diff estimator until shown their current one is broken — Goodman-Bacon's 2021 result that two-way fixed effects (TWFE) is biased even under parallel trends is what made differential-timing estimators stick, and the same wedge motivates learning the continuous-treatment paper by Callaway, Goodman-Bacon and Sant'Anna ("CBS," conditionally accepted at AER). The strategy is "backwards engineering" (per Sant'Anna): run the regression, then crack the coefficient open via Frisch-Waugh-Lovell to see what estimand it actually targets — chosen over "forwards engineering" because people must first learn what the coefficient they love means. The application is Lu and Yu (2015, AEJ:Applied) on China's 2001 WTO accession, using predicted tariff cuts as a continuous dose to show larger cuts reduced within-industry markup dispersion. The build uses Claude Code skills: /split-pdf (chunk papers, summarize each, then summarize the whole), /beautiful_deck (Beamer slides on Aristotle's ethos/pathos/logos, one idea per slide), /tikz (fix label/arrow collisions via Bezier depth formulas), and /referee2 (fresh-terminal adversarial deck audit). Part 1 only builds architecture and the deck.

continuous diff-in-diffTWFEBacon decompositionClaude Coderesearch workflow

Decomposing the TWFE regression coefficient with continuous treatment dosage using FWL

TIER 4 Apr 15, 2026

A two-period TWFE regression of outcomes on unit/time fixed effects and a continuous dose (e.g. how much a municipality raises the minimum wage, not just whether) reduces, via Frisch-Waugh-Lovell, to the OLS slope of the unit-level first difference on dose. Working the "Levels" row of Callaway, Goodman-Bacon and Sant'Anna's Table 1: condition on D, split its mass at zero from its positive support, add and subtract m(0). The m(0) terms cancel because mean deviations sum to zero, leaving the coefficient as an integration-weighted average of dose-level calculations—still purely algebraic, no causality yet.

continuous diff-in-diffTWFEFrisch-Waugh-Lovellderivationdecomposition weights

Making a shiny to illustrate the TWFE continuous weights

TIER 4 Apr 20, 2026

Callaway, Goodman-Bacon and Sant'Anna's continuous-dose diff-in-diff paper decomposes the TWFE coefficient via Frisch-Waugh-Lovell; the "level" weight has three ingredients — mean dose E[D]=0.164, variance 0.0202 (denominator), and density f_D(l), integrated over doses. The decisive feature is a sign flip: the weight is zero exactly at the mean, positive for above-average doses, negative below. A Claude Code-built shiny app sliders the dose to show this; a Gaussian kernel spuriously smeared density left of the smallest observed dose, fixed by adding min/max dashed lines.

continuous diff-in-diffTWFER Shinydecomposition weightsClaude Code

Claude Code Series (part 8): Resurrecting and Extending an Old Abortion Paper Towards Using Continuous Diff-in-Diff

TIER 4 Jan 20, 2026

Reviving a 2019 Journal of Human Resources paper (Cunningham, Schlosser, Lindo, Myers) on Texas HB2 abortion-clinic closures, using Claude Code to re-estimate distance effects under Callaway, Goodman-Bacon, and Sant'Anna's conditionally-accepted AER continuous diff-in-diff method. Claude audited two rival distance datasets: an earlier thesis version backdated 2010 distances to 2006 (assuming no pre-period closures, killing within-county variation) versus Caitlin Myers's hand-tracked "ground truth." They diverge 8% of the time by missing New Mexico/Oklahoma clinics—Lubbock reads 307 miles versus 78. Beamer/TikZ decks, CLAUDE.md, todo.md, and logs serve as memory. Unsolved: urban counties (42% of Texas) supply no treatment variation, breaking the counterfactual.

claude-codecontinuous-diff-in-diffabortion-accessidentificationdata-audit

Claude Code Series (part 8): Resurrecting and Extending an Old Abortion Paper Towards Using Continuous Diff-in-Diff [original send]

TIER 4 Jan 20, 2026

Identical content to 0111: a Claude Code case study reviving the JHR Texas HB2 abortion project for continuous diff-in-diff, with a thesis-vs-JHR distance data audit and the urban-counterfactual identification problem. This is the original post that was later deleted and reposted as 0111 because it had the wrong video attached; substantively it is the same essay.

claude-codecontinuous-diff-in-diffabortion-accessidentificationdata-audit

Vertical regression and selection bias in diff-in-diff (plus some pictures of Pisa and Stresa Italy)

TIER 5 Jun 12, 2026

Parallel-trends is just selection bias on a derivative: where simple comparisons carry selection bias as the gap in mean untreated potential outcome Y(0) between treatment and comparison groups, diff-in-diff's "non-parallel trends bias" is the same gap in first differences of E[Y(0)]. Working the 2x2 in four moves (expectations, swap realized for potential outcomes under no-anticipation, add a zero, rearrange) isolates the ATT plus that lone bias term — proof the calculation isn't diff-in-diff without the assumption. The bias term is itself a 2x2 on Y(0). Following Imbens, the horizontal (first-differences) and vertical (within-time group-difference) regressions are algebraically identical; TWFE does both. Cites the JEL Practitioner's Guide and Mostly Harmless Econometrics.

difference-in-differencesparallel trendsselection biassynthetic controlvertical regression

Diff-in-diff foundations — parallel trends, selection bias, and covariates

1 tier-5 · 1 tier-4

The bedrock tutorials on what difference-in-differences actually assumes. Cunningham keeps reframing parallel trends as a selection-bias condition stated in first-differenced untreated potential outcomes, shows the 2x2 estimator computed many numerically-identical ways to demystify the algebra, and works through when covariate balance does and doesn't matter for the estimand. (Note: issue 0002, which most fully unifies parallel trends with selection bias and vertical regression, is filed in Theme 4 alongside the continuous-DiD arc it extends.)

Diff-in-diff can be written down six ways!

TIER 5 Jun 4, 2026

A simple 2x2 diff-in-diff is "four averages and three subtractions," and the identical point estimate can be reached six ways. Two manual orderings: first-difference each group (after minus before, treatment first) then subtract, or take group differences (treatment minus control) each period then subtract; reordering works because subtraction signs travel. Four OLS specs match it numerically: a saturated treatment-times-post dummy regression, two-way fixed effects, regressing first-differenced outcomes on a treatment dummy, and regressing group differences on a post dummy. Demonstrated with Card-Krueger and castle-doctrine Stata code; diff-in-diff isn't one regression.

difference-in-differences2x2TWFEStata coderegression equivalence

Should I Include Covariates in Diff-in-Diff?

TIER 4 Jun 1, 2026

The belief that diff-in-diff estimates shifting when covariates are added discredits them is wrong. Under parallel trends you need no controls. Worked example: college-vs-high-school earnings where male trends grow +10, female +8. First-differencing wipes out level effects (alpha) but not sex-driven trend differences. When groups are balanced (75% male in both), both grow 9.5, so unconditional parallel trends holds and the 2x2 unbiasedly estimates the ATT without controlling for sex.

difference-in-differencescovariatesconditional parallel trendscovariate balanceATT

Callaway–Sant'Anna and the perils of staggered DiD with covariates

3 tier-5 · 4 tier-4

A sharp, practitioner-facing cluster on a single underappreciated fact: Callaway–Sant'Anna with covariates secretly fits one propensity-score logit *per treatment cohort*, and that hidden stage is where things break. Cunningham develops the consequences across several posts — Peduzzi's events-per-variable rule means the binding constraint is treated units per cohort (so U.S. state-level staggered panels with singleton treated states routinely violate it), some packages quietly drop covariates or fail to converge, and an apple-to-apple audit running identical specs across six R/Stata/Python packages produced ATT estimates ranging up to ~5x apart, driven by covariate handling, matrix conditioning, and near-separation in the logit. The constructive payoff: zero-covariate baselines, z-scoring covariates, reporting package and version, and switching to regression adjustment (which has no events-per-variable problem).

The Many Logits of Callaway and Sant'Anna and Why It Matters for Your Covariates

TIER 5 Apr 1, 2026

The Callaway and Sant'Anna staggered-adoption diff-in-diff estimator (11,000+ cites, the default AI agents pick) cannot absorb as many covariates as users assume, because its IPW version estimates one logit per treatment cohort, not one regression overall. Covariates enter through propensity scores fit separately for each of the (say) 14 treatment years; what constrains each fit is treated units in that cohort, not the total. Peduzzi et al. (1996)'s "ten events per variable" rule then bites locally: a 2003 cohort with 47 treated units chokes on covariates a 2006 cohort with 216 handles fine, risking flat likelihoods, overfitting, and perfect separation. Worse, some packages silently drop all covariates and regress treatment on a constant. Run the 14 logits yourself; state panels with 50 units almost always trigger this. Regression adjustment sidesteps logits but trades one problem for another.

Callaway-Sant'Annadiff-in-diffpropensity scorecovariatesevents per variable

Claude Code 31: Apple-to-Apple Audit of Six Callaway and Sant'Anna packages

TIER 5 Mar 12, 2026

Six independent software packages implementing the identical Callaway and Sant'Anna (2021) estimator on identical data and specifications return wildly different answers — ATTs from 0.00 to 2.38 homicides per capita on a mental-health-closure question, one package saying half a standard deviation, another more than double. Claude Code wrote ~96 scripts (16 covariate specifications across did, ddml, csdid, csdid2, differences, diff-diff in R/Stata/Python) on Brazilian municipal data, 2002-2016. With zero covariates all packages agreed to four decimals (ATT ≈ 0.31); adding one variable, poptotaltrend (population × year, reaching ~10 billion), fanned estimates out. A two-way ANOVA: 40% of variation from specification, 16% from package, 44% from interaction. Two culprits: floating-point matrix-inversion failure from huge condition numbers (fixable by z-scoring), and near-separation in the propensity score — outlier structure scaling doesn't fix, since each package's logit optimizer (MLE vs. csdid's inverse-probability tilting) handles perfect prediction differently. R's did silently returned zeros. This is undocumented publication bias: package choice is a substantive decision no journal requires reporting. Recommendations: report package/version, standardize covariates, run a zero-covariate baseline, and use AI agents for such cross-language code audits.

callaway-santannapackage-variationpropensity-scorenear-separationcode-audit

Claude Code 53: Applied econometrics will require a detailed checklist

TIER 5 Jun 8, 2026

Automating applied econometrics raises, not lowers, the returns to econometrics knowledge, because automation replaces production but not verification. Letting Claude "rip" with /skills backfires: it specification-searches, silently writes up whichever method fits the prompt, and buries decisions in a JSON. A four-hour Callaway-Sant'Anna (CSDID) debugging ordeal—NAs in 2x2s—turned out to be the CRAN `did` build versus the newer GitHub one; two Claudes confidently ran down wrong conjectures throughout. Agents make reasoning mistakes, unlike human coding mistakes, and strip away the "epistemological feeling" of knowing what produced what. The fix is structured analog workflows—checklists (Gawande's Checklist Manifesto; "Pedro's Checklist"): name the estimand in potential outcomes, tabulate treated units per cohort (7-10 events per covariate for the propensity score), pre-register aggregation, plot rollouts. Constraint: zero error.

AI for researchCallaway-Sant'Annastaggered diff-in-diffchecklist/workflowestimand specification

Claude Code 34: Using "Dispatch" on my phone with Claude Code to revisit the cannabis paper

TIER 4 Mar 20, 2026

Framed around the new phone-based "Dispatch" remote-control feature, but the substantive core is a methods lesson: Callaway-Sant'Anna with covariates secretly fits a per-cohort logit propensity score, and Peduzzi et al.'s "events per variable" (EPV) rule means you need ~10 treated units per covariate per cohort-year to avoid biased propensity scores. Cunningham argues U.S. state-level staggered DiD routinely violates this (singleton treated states), and the fixes are dropping to county-level data or switching from doubly-robust/IPW to regression adjustment, which has no events-per-variable problem. Useful for anyone running CS with covariates.

callaway-santannaevents-per-variablepropensity-scoreconditional-parallel-trendsclaude-code

Claude Code 24: Multiple Agents Auditing Your Diff-in-Diff Code (Part 1)

TIER 4 Feb 25, 2026

Treat LLM coding hallucinations as classical measurement error: random, language-specific syntax slips that systematically corrupt downstream results. A real example—Stata's `replace olddog = 10 if olddog>10` silently recodes missing values too unless `& olddog~=.` is appended—shows how one bad random draw cascades through a pipeline unnoticed since the code still runs. The fix exploits independence: if errors across R, Python, and Stata are uncorrelated (Cov=0), the chance all three hallucinate identically is the product of tiny probabilities. So have Claude Code aggressively audit code like a health inspector AND fully replicate the pipeline in two other languages, demanding identical tables to several digits. This works for deterministic methods (OLS, diff-in-diff, IV) but not random ones (bootstrap, MCMC, ML). Case study: five diff-in-diff packages (csdid, csdid2, did, differences, diff-diff) on a Brazilian CAPS deinstitutionalization study.

code auditdifference-in-differencesmeasurement errorLLM hallucinationreplication

Claude Code 26: Multiple Agents Auditing Your Callaway and Sant'Anna Diff-in-Diff (Part 2)

TIER 4 Feb 27, 2026

Hand the same DiD problem to fifteen isolated AI agents and the spread of estimates becomes a cheap robustness audit — an automated version of the "multi-analyst" findings (Silberzahn's 29 teams; Huntington-Klein's seven economists, whose cross-analyst SD ran 3-4x the typical standard error; Menkveld's "non-standard errors"; Borjas-Breznau, where ideology predicted the effect's sign). The setup: 15 fresh `claude -p` agents, three each across five Callaway-Sant'Anna packages (Python's differences/diff-diff, R's did, Stata's csdid/csdid2), all run on Dias-Fontes's Brazilian CAPS mental-health rollout over 5,476 municipalities, under strict isolation. Result: every structural choice was unanimous (not-yet-treated controls, universal base period, no trimming, 15/15). Variation lived entirely in covariate selection — log GDP near-universal, geographics rejected, but the confounder-versus-mediator line (poverty 10/15, health spending 7/15, Bolsa Familia 2/15) shifted agent to agent. Whether that disagreement actually moves the ATT is held for Part 2.

difference-in-differencesCallaway-Sant'Annanon-standard errorsmulti-analyst designAI agents

Claude Code 28: Multiple Agents Auditing Your Callaway and Sant'Anna Diff-in-Diff (Part 3)

TIER 4 Mar 4, 2026

Covariate selection, not the choice of estimator, drives most variation in Callaway-Sant'Anna difference-in-differences estimates. Across 15 runs on one fixed dataset through five language-packages (Python's differences and diff-diff, R's did, Stata's csdid and csdid2), all positive, the average ATT sat near 0.4, but an R agent using only state fixed effects got 0.17 while a diff-diff agent with 8 covariates got 1.83 — nearly 11 times larger, each rationale defensible. Roughly 77% of variation came from between-package differences. The "non-standard error" punchline: the standard deviation across the 15 point estimates (0.442) is 2.4 times the average reported standard error (0.185) — sampling-based standard errors capture none of the team/discretion uncertainty. Picking g-5 over g-1 as baseline silently changes the parallel-trends assumption (always estimated as long differences). Next: doubleml added, 20 agents per package for 120 estimates, free covariate range.

callaway-santannamany-analyst-designcovariate-selectionnon-standard-errorscode-audit

Synthetic control, matching, and identification theory

2 tier-5 · 0 tier-4

The two most reference-worthy pure-method tutorials in the archive, both about identification beyond the DiD parallel-trends frame. One derives the identifying assumption of synthetic control under the Abadie–Diamond–Hainmueller factor model and contrasts outcome-model identification with design-based (random-assignment) identification, with Monte Carlo code showing why a long, well-fitting pre-period guards against matching on noise. The other shows the Abadie–Imbens bias correction for nearest-neighbor matching written two numerically-identical ways — standard imputation and an "augmentation" form that slides the matched control's outcome along the regression line — and connects it to augmented synthetic control and the Wald/2SLS equivalence. Together they form Cunningham's clearest writing on what makes a non-experimental estimate credible.

Identification in Synthetic Control

TIER 5 Dec 12, 2025

Synthetic control's identifying assumption, from Abadie, Diamond and Hainmueller (2010), is an outcome-model claim: untreated outcomes Y(0) follow a factor model — constant α, observable X with time-varying loading β, and crucially an unobserved unit factor μ (e.g. CEO ability) with its own time-varying loading, plus transitory shocks ε. This places synth alongside difference-in-differences (which characterizes potential outcomes), not the design-based tradition (which characterizes treatment assignment). Under random assignment the structure cancels in expectation, any weights work, and you needn't invoke the factor model at all. Without randomization the donor weights matter: matching well on observable X (ideally lagged outcomes) zeroes the X residual, but bias remains from the unobserved μ and transitory ε. ADH's argument is that a long pre-period plus good pre-fit forces μ to be approximately matched too — otherwise the fit relies on noise (overfitting). Abadie's 2021 JEL primer frames bias control around transitory-shock scale and pre-period length T₀. Synth was born modeling Basque ETA terrorism.

synthetic controlfactor modelidentificationdiff-in-diffcausal inference

Two Ways to See Abadie-Imbens Bias Correction (And Why It Might Matter)

TIER 5 Dec 24, 2025

The Abadie-Imbens (2011) bias correction for nearest-neighbor matching has two numerically identical forms. The standard imputation form regresses Y on X using matched controls only, predicts counterfactual outcomes μ̂⁰ at the treated unit and its match, and subtracts that difference from each naive treatment effect — fixing the non-√N bias caused by inexact matches on continuous covariates. The equivalent augmentation form ignores predicted outcomes: it takes the match's actual outcome and slides it along the regression line by (slope × covariate discrepancy ΔX). Algebraically the intercepts cancel when differencing the imputed outcomes, leaving only the slope term — proving equivalence. Ben-Michael, Feller, Rothstein (2021) use exactly this augmentation in augmented synthetic control, citing Abadie-Imbens. Parallel to how the Wald IV estimator equals 2SLS, the augmentation lens (focused on X, not Y) may make otherwise-opaque fixes visible.

matchingAbadie-Imbensbias correctionaugmented synthetic controlregression adjustment

LLMs as measurement instruments — text-as-data and replication

1 tier-5 · 1 tier-4

A focused empirical arc treating large language models not as writing assistants but as classification and measurement tools, and asking whether they reproduce trained-classifier results. Cunningham re-classifies hundreds of thousands of congressional speeches from Card et al.'s PNAS immigration-rhetoric paper with gpt-4o-mini via the Batch API ($11, ~2.6 hours), gets 69% agreement with a fine-tuned RoBERTa, and shows the paper's polarization findings survive — because disagreements cluster at the decision boundary and largely cancel in the net-tone difference. The capstone finding (the gpt-4o-mini thermometer-heaping result, issue 0088) and the setup half of the replication (issue 0101) are filed in Themes 2 and 3 respectively; the lower-tier installments of the "why does it work" sub-arc (0094, 0095) are counted but not keepers for this guide. The cluster also includes the Ludwig–Mullainathan–Rambachan framework distinguishing prediction tasks (which require no training leakage) from estimation tasks (which need a small human-coded validation sample to debias LLM labels).

Claude Code 15: The Results Are In: Can LLMs Replicate a PNAS Paper? (Part 2)

TIER 5 Feb 6, 2026

A zero-shot gpt-4o-mini run classified all 304,995 speeches from Card et al.'s 2022 PNAS study of 140 years of US immigration rhetoric, agreeing with the original fine-tuned RoBERTa classifier on 69% of speeches—for $10.99 and 2.6 hours of batch processing, versus the weeks of human annotation (7,626 labels) RoBERTa required. 69% beats chance (~33% for three classes) and sits near the moderate human inter-annotator agreement (Krippendorff's α=0.48). Disagreements are systematic but benign: the LLM hedges toward NEUTRAL when uncertain (diagonal agreement 85% neutral, 63% pro, 51% anti), while direct polarity flips stay rare (PRO→ANTI 3.7%, ANTI→PRO 4.9%)—it disputes opinionated-vs-neutral, not direction. The substantive findings survive: post-1970s partisan polarization, country-of-origin ordering (Italy>China>Mexico), and Texas patterns all replicate, with more volatility. Conclusion: cheap general LLMs plus Claude Code make exploratory text classification accessible without fine-tuning, GPUs, or annotation budgets. Process caveat: an independent "Referee 2" review caught real bugs, but running it inside the same context window defeats the purpose.

llm-classificationreplicationbatch-apitext-analysisai-for-research

Explainer of Ludwig, Mullainathan and Rambachan's 2026 Econometrics of LLM Paper

TIER 4 Mar 25, 2026

High LLM-labeling accuracy does not protect regression estimates: errors can correlate with covariates and destroy inference. Ludwig, Mullainathan, and Rambachan split LLM uses into prediction (forecasting an outcome) and estimation (measuring a concept fed downstream into a regression), each needing a different discipline. Prediction requires "no training leakage" — prompt tricks like "ignore data past this date" fail; GPT-4o reproduced 344 of 10,000 Congressional bill descriptions verbatim. Estimation requires a small, human-coded validation subsample used not to replace LLM labels but to debias them — echoing double/debiased ML.

LLM econometricsmeasurement errorvalidation sampletext classificationprediction vs estimation

Teaching econometrics in the age of AI

0 tier-5 · 2 tier-4

Cunningham's pedagogy essays, unified by one claim: generative AI has decoupled task-completion from learning, which historically were the same act, so teaching must change to force the time-on-mechanics that builds real human capital. The concrete moves are by-hand spreadsheet worksheets (decomposing a difference in means into ATE + selection bias + reweighted ATT/ATU; computing regression coefficients manually and checking against R), and — paradoxically — using AI to *manufacture* teaching examples: having Claude generate exemplar papers in three genres (descriptive, predictive, causal) so students can study the rhetoric of each, even while he bans AI in his courses. The "rules before strategy" framing argues a skill splits into two halves that should be taught separately.

Rules before strategy or how I'm trying to teach statistics and causal inference

TIER 4 Nov 13, 2025

Rules and strategy are different kinds of human capital and should be taught separately: teach a new card player only the mechanics, let them play, and strategy emerges on its own (Kahneman's System 1 versus System 2). Applied to Gov 50, a 200-student Harvard methods course split across young/old and Gov/non-Gov concentrators, this means brute-force drilling of mechanics. With generative AI, completing a task and learning from it have decoupled for the first time — output and human capital used to be produced together; a novice leaning entirely on AI spends zero time on the activity that actually builds understanding. So students are "paid" via extra-credit spreadsheet worksheets to compute things by hand: the decomposition formula (simple mean comparison = ATE + selection bias + reweighted ATT/ATU difference), bivariate regression coefficients as "scaled covariance" (covariance over variance) on 5 observations, and two-covariate multivariate coefficients on 4 — each verified against R's lm(y~x). Confusion then reveals not bad math but not grasping the rules. Pedagogy is endogenous to course goals: reasoning quantitatively requires underlying theory, not button-pushing. Scaling efforts (office hours, research labs) mostly didn't, but students report learning otherwise-elusive material.

teachingpedagogyregression mechanicsai and learningcausal inference

A professor's use case for AI generated papers

TIER 4 Apr 16, 2026

AI agents now write journal-submittable empirical papers end to end, flattening the human-versus-machine isoquant for cognitive output. Cunningham still bans AI in his PhD and undergrad stats courses, betting that 10-20 hour problem sets, frustration, and failure are how learning happens. Yet he found a use case: students needed example papers across his three statistical genres (description, prediction, causal inference), and while he can model causal papers, he can't articulate the skeleton of descriptive or predictive ones. So Claude Code generated three unsupervised exemplars (polarization measurement, COMPAS recidivism prediction, an Acemoglu-Johnson-Robinson settler-mortality IV) showing each genre's distinct rhetoric and exhibits.

AI-generated paperseconometrics teachingresearch genresClaude Codeacademic policy

Applied empirical essays and the sociology of the discipline

0 tier-5 · 5 tier-4

The remaining substantive essays that sit outside the methods tutorials and the AI-workflow spine: applied economics, the statistics of research integrity, and the intellectual history of the field. Cunningham dispatches Claude Code to crawl EDGAR and build an HHI analysis of Match Group's dating-app portfolio (a zero-price market that evades antitrust); works through why reconstructing t-statistics from rounded coefficients manufactures false bunching at t=2 (and publicly retracts his own p-hacking claim when he finds the artifact); distinguishes three sources of uncertainty — sampling, design-based, and analyst/researcher uncertainty that standard errors never capture, now operationalizable via many-analyst automation; and maps the family tree of causal inference through Orley Ashenfelter's academic lineage.

The Orley Genealogy Project: Mapping the Family Tree of Causal Inference

TIER 4 Dec 31, 2025

Modern causal inference in economics descends from Orley Ashenfelter and Princeton's Industrial Relations Section, a quantitative-labor lineage that merged with Don Rubin's potential-outcomes framework once Imbens and Angrist linked instrumental variables to the Rubin model. Three streams—Princeton's natural experiments, Harvard's experimental design, Chicago's Heckman/structural—form "two rivers converging." A genealogy database (~1,100 economists across four generations: Card, Heckman, Angrist, Duflo, Abadie) crowdsources missing advisor-student edges, source-verified, treating people as the real confounders.

causal inference historyOrley Ashenfelteracademic genealogyPrinceton IRScredibility revolution

Swiping Under Monopoly: Market Power and Welfare in Online Dating

TIER 4 Apr 6, 2026

Match Group owns Tinder, Hinge, OkCupid, Plenty of Fish and more, controlling roughly 65% of a market where over half of couples now meet online; a ban from one app ejects you from all. Revenue-share HHI runs 4,600–5,600 versus the 2,500 "highly concentrated" threshold, yet DOJ ignores it because antitrust's price-focused consumer-welfare framework can't see zero-priced products. Match acquired 25+ firms with essentially no merger review—a regulatory failure if dating is critical social infrastructure.

antitrustmarket concentrationzero-price marketsonline datingClaude Code research

Claude Code 37: Building an Understanding of Rounding for the Purposes of Publishing and Evidence for P-Hacking

TIER 4 Apr 2, 2026

Extracting coefficients and standard errors from published tables to reconstruct t-statistics produces false p-hacking evidence, because the rounding done purely for display collapses the continuous t-stat ratio into discrete heaps near 2. Authors' actual hypothesis tests use unrounded numbers, so their p-values are unaffected; only third-party reconstruction inherits the artifact. Ratios of exactly 2 are the second-most-common single-digit-integer ratio (after 1), so small-scale outcomes—coefficients like 0.01527 with leading zeros—heap there once rounded; large-scale outcomes (income on a college dummy) don't, since rounding only touches fractional digits and t-statistics are scale-invariant. Brodeur et al.'s diff-in-differences p-hacking signal vanished after a corrected fix (IV survived); the same flaw afflicted Cunningham's DiD-heavy APE-project analysis.

p-hackingrounding artifactst-statisticsShiny appresearch methods

Claude Code 36: I Was Wrong About P-Hacking (And Here's What I Actually Found)

TIER 4 Mar 30, 2026

A t-statistic spike near 1.96 in 651 AI-written economics papers wasn't p-hacking; it was an artifact of dividing rounded inputs. Because papers report coefficients and standard errors rounded (e.g., 0.035/0.021 = 1.67, but 0.04/0.02 = 2.0), reconstructing t-stats from those LaTeX-table values collapses many "near-2" cases onto exactly 2.0. Small numbers dominate in labor economics (log outcomes, proportions, linear-probability coefficients), and 2:1 is the simplest small-integer ratio, so the heap lands at 2, not 1.96. A simulation of smooth, unmanipulated data still piled 7.2% of t-stats onto 2.0. Brodeur et al. (2020) avoided this by using software-output t-stats, not reconstructions. A donut-hole drop of exact-2 cases flattens the reported Brodeur ratio from 1.52 to 1.02—no bunching.

p-hackingrounding artifactscorrectionBrodeur testAI-generated papers

If Non-Standard Errors Are Measuring Real Uncertainty, Should We Report Them?

TIER 4 Mar 6, 2026

A third source of statistical uncertainty exists that standard errors never capture: variation from who the researcher is. Sampling inference holds the researcher fixed and resamples the population; design-based inference (Fisher's lady tasting tea, randomization inference) holds the sample fixed and permutes treatment assignment. Both ignore that the analyst is not a transparent pipe. Silberzahn's many-analyst study sent one dataset and question to 29 teams and got a spread of estimates from defensible but subjective choices alone. A pipeline with ten decision points, three reasonable options each, yields 3^10 ≈ 59,049 possible estimates, none normally reported. Claude Code makes this operational: running the same Callaway–Sant'Anna diff-in-diff across five packages and varying covariates showed 75% of total variation traced to whether Python, Stata, or R was used. The hard unsolved part is automatically enumerating the discretionary (endogenous) nodes where reasonable analysts diverge, then perturbing them to build intervals and a p-value—how often is a node pivotal?

non-standard-errorsmany-analyst-designresearcher-uncertaintyinferenceclaude-code