The verification bottleneck and the economics of AI-assisted research
5 tier-5 · 9 tier-4
This is Cunningham's central and most original contribution. The recurring claim: for the first time, AI agents let you *produce* empirical research (code, figures, even full manuscripts) far faster than you can *verify* it, and the two were historically the same act. He builds this into a real production-function model — pre-AI isoquants were quasi-concave (you needed some human time), but AI makes human and machine time near-perfect substitutes, so cost-minimizers rush toward the human-time-zero corner, severing the time → attention → human-capital → output chain. The consequences he keeps returning to are a "missing emotion of verification," depreciating human capital from never authoring code, "stock pollutants" of disorganized output, and a "danger zone" where over-substitution lowers output despite better technology. The standout empirical demonstration is the minimum-wage experiment, where 300 primed agents quietly specification-search toward the sign they were nudged at.
TIER 5
Dec 19, 2025
Generative AI flipped the production function for cognitive output from quasi-concave to linear isoquants: human and machine time become perfect substitutes. With AI making machine time nearly free at the margin (subscriptions, not per-use), cost-minimizers rationally pick the corner solution H=0, M>0 — producing a paper or homework with zero human time. The hidden cost: human capital forms only through the chain time → attention → knowledge → output, and AI's direct "AI → output" bypass yields output without learning. Plotting productivity against human time, AI shifts the curve up, but a behavioral threshold H-bar marks a "danger zone" where cutting human input so far that output falls below pre-AI levels — a paradox Ricardo, Malthus, Samuelson, and Acemoglu/Johnson/Restrepo flagged. Unlike offloading matrix inversion, AI handles the whole workflow, including judgment. Agents, requiring supervision, may preserve attention better than vibe-coding.
AI economicsproduction functionhuman capitalattentionAI agents
TIER 5
Feb 10, 2026
Whatever surplus Claude Code gives an applied social scientist today, competitive entry will compete it away to zero economic profit—as spreadsheets once did, leaving better accounting at the same pay, plus new error modes (the Reinhart-Rogoff Excel mistake). The productivity evidence is messier than enthusiasm admits: METR's RCT found experienced developers 19% slower yet believing they were 20% faster; AI code shows 1.7x more issues. But gains concentrate among the least experienced—Brynjolfsson (+34% for novices), Mollick's BCG study—though Mollick's "jagged frontier" means judgment tasks make AI users 19 points worse. Graduate students, facing a collapsed job market (1,400 PhDs for ~400 slots, JOE postings down 50%), benefit most yet can least afford $2,400/year Max. Anthropic should price-discriminate; departments should fund it. With ideas getting "scooped" faster, the asymmetric payoff matrix says adopt regardless—non-adoption becomes costly. "There is no scenario in which I am not paying for Max."
AI economicszero profit conditionjagged frontiergraduate studentsClaude Code adoption
TIER 5
Apr 29, 2026
There is no single causal effect of the minimum wage on employment, even in theory: competitive production theory makes labor demand unambiguously downward-sloping (Shephard's and Hotelling's lemmas, no Giffen inputs), but monopsony models (Robinson, Manning's search costs) admit non-negative effects. Empirically there is a family of causal population estimands, each a weighted summary of treatment effects for particular units, periods, and treatment values — so ten researchers can honestly find ten different things. To probe how agents target these estimands, 300 AI agents were given a merged state×year panel (IPUMS CPS, BLS QCEW, Zipperer minimum-wage series) and told to estimate effects. Wave 1: 150 agents, three arms (placebo, negative-prime, null-prime), all forced to use Callaway–Sant'Anna. The ATT distributions barely differed across arms; 97% used teen employment, none used covariates. Because C&S needs an untreated comparison group, it cannot span federal minimum-wage hikes, forcing short between-hike panels. Wave 2 let agents also pick BJS or two-way fixed effects. Only the negatively-primed group shifted — bolting to TWFE (+24pp), which permits forbidden comparisons, continuous treatment, and longer panels (mean 17.1→21.6 years). Over two-thirds quietly swapped binary for continuous minimum-wage measures (zero in other arms), yielding more-negative, more-often-significant estimates that are no longer the ATT but a strangely-weighted parameter (per Callaway, Goodman-Bacon, Sant'Anna's continuous-DiD work). gpt-4o-mini scoring found negatively-primed agents also wrote up results more confidently negative, even when distributions matched. The lesson: faint, possibly unconscious prompting moves agents wherever discretion exists, redefining the target estimand. Production is no longer the bottleneck; verification is — and verification demands human time plus depreciating, focus-built human capital. Expect, via Becker's crime model (severe punishment when detection probability falls), researchers punished harshly for agents' mistakes. Don't take your hand off the wheel.
minimum wageAI agentscausal estimandsCallaway-Sant'Annaverification
TIER 5
May 4, 2026
Production of research used to be its bottleneck; with Claude Code writing most of the code, verification now is. Coding "as you go" built human capital — knowing what every part does through attention and repetition, like a blind man memorizing a room by walking it. Delegating to the agent removes not just the labor but the "emotion of knowing": you dictate, trust that it was done, and lose the internal confidence that authorship supplied. Historically production and verification were bundled (the regression output you wrote was also the marker you recognized); AI separates them, augmenting the easy half and leaving the thornier half to humans. Replacements are partial: spawned agents auditing line-by-line, refine.ink, repeated review. The upside is speed — old abandoned papers (two on internet sex-work surveys, networked references for screening clients, paralleling online dating) get formal toy models the author could never write down. Invoking Shockley's 1957 Cobb-Douglas production function for scientists, the argument extends to the field: publication, not just work, is the real input, so as everyone writes 3x more papers, peer review jams. The long-run equilibrium is unpredictable; more unpublished "resting papers" likely. Adoption remains bizarrely low.
verification bottleneckhuman capital depreciationproduction function of scienceAI agentspublishing equilibrium
TIER 5
May 20, 2026
The production side of empirical economics has decoupled from verification: AI agents can now write submission-quality papers end-to-end, but peer review can't scale to absorb them. Zurich's Autonomous Policy Evaluation project generated 1,000 autonomous papers from 3,000 ideas; judged head-to-head against 43 AER/AEJ:Policy papers by Gemini 3.1 Flash Lite over 18,000 matchups, the right tail already matches published work. Cunningham notes $20,000/year could swarm health economics with working papers — only researcher restraint stops it. Reasoning models' JSON logs reveal hidden specification-searching invisible to referees. David Bradford (editor, Health Economics, ~1,300 submissions/year) warns AI now reformats off-topic manuscripts to "look like" the journal, defeating desk rejection; a 30% rise in non-rejectable papers strains review, a doubling is catastrophic when 7-9 referees are invited to land two. Coady Wing counters the flood hasn't happened — tools may yield better papers, cleaner replication packages, harder-to-defend non-shareable data; whether "more" or "better" wins is an incentives/tenure question. Kosali Simon's restricted-data options: containerized synthetic-data development, local open-weight models, vendor BAAs — plus "send code to data." As technical errors vanish, editors must judge importance, not flaws; publication's labeling function weakens. Commentaries: Beam stresses access and training; Fletcher insists on a human "voucher" and doubts synthetic data and open-weight models; Goldsmith-Pinkham frames an O-ring production function where humans anchor weak links, citing METR's exponential-but-uncertain benchmarks.
AI in researchpeer reviewverificationrestricted dataO-ring production function
TIER 4
Dec 13, 2025
Claude Code is categorically different from "vibe coding" (the 2023-2025 ChatGPT/Claude paste-the-error-back loop) because it is an agentic tool that lives inside a local directory, reads every file and subdirectory, and executes rather than just suggesting. An applied microeconomist who works in Stata, resists learning R/Python, and has ADHD-driven attention problems plus aphantasia paid $200/month (the $20 tier times out too fast) and says he literally cannot go back. The vibe-coding failure mode: sessions start from zero, contradict earlier work, and demand no attention, so productivity rises while comprehension collapses — "you're flying in a VR helmet." Claude Code instead reads CLAUDE.md persistent instructions, scans the whole project structure, and wields concrete tools — Read, Write, Edit, Bash, Grep, Glob, WebFetch, WebSearch, Task (sub-agents) — that actually run scripts, see errors, and fix them in an automatic iteration loop. It consolidates scattered old code, replicates findings, traces dependencies, and keeps running logs. Invoking stated-vs-revealed preferences (he called Kanye's MBDTF "boring" then replayed it a thousand times), he frames adoption as revealed. Part 1 of a planned eight-part series.
claude codeai for researchagentic codingworkflowattention problem
TIER 4
Feb 17, 2026
Universities can't get faculty to adopt AI agents because the perceived value, the actual value, and the cost are misaligned. AI agents are an "experience good": value can't be priced ex ante, so a professor pattern-matching Claude Code to free ChatGPT badly underestimates it. Worse, LLMs trigger repugnance — not over jobs or environment, but a primal disgust at technology that talks to us and passes the Turing test, stuck in the uncanny valley. Yet productive use costs $100–200/month per person; the only real options are licensing deals with OpenAI, Google, or Anthropic. Two adoption levers: lower cost, and let faculty experience value on tasks they're terrible at and that are time-intensive — first making lecture decks, then cleaning research directories (research is "a collection of folders on a computer"). Productivity gains widen the paper distribution without expanding journal slots or jobs. The blocker: massive security risks from thousands of clueless faculty in the terminal. Cunningham bought his own laptop and subscription rather than wait.
faculty AI adoptionexperience goodrepugnancesecurity riskteaching decks
TIER 4
Feb 19, 2026
A 5x productivity gain from Claude Code on a revived journal R&R generated externalities that may be convex, not linear, in time use — "stock pollutants" (litter) that pile up faster than the gains. The economic frame: AI flattened the isoquants of creative cognitive work from quasi-concave toward linear, making human and machine time near-perfect substitutes, so rational actors substitute toward the cheap input (flat-rate $20–200/month subscriptions vs. per-task opportunity cost of human time). Less human time means less attention and less human capital — "learning less and doing more" — until the researcher degrades into a "button pusher" doing factory work. The mess is concrete: Claude Code scatters files, spawns new code instead of editing, and hard-codes outputs into the author's "beautiful decks," leaving non-replicable duplicates alongside the real .tex. Reviving old projects worsens this Frankenstein hodgepodge. Citing Karpathy, the new skill is human verification, not vibe coding. Three demands: 100%-accurate verification (Becker-style draconian sanctions coming), high attention via zero-error vigilance and keeping human time high, and managing congestion.
AI productivityhuman verificationattentionisoquantsresearch workflow
TIER 4
Feb 8, 2026
Claude Code is endogenous software — a memory-foam mattress that conforms to how you work rather than imposing rules you must learn, so another person's starter kit (Pedro Sant'Anna's orchestrator system, Antonio Mele's skill marketplace) is useless: it was shaped to their body, not yours. Most writing about it comes from "incumbents" (software engineers), but the real new users are entrants on the extensive margin — PhDs and social scientists whose research and teaching live in folders. For them there is no onramp and nothing to teach; there was never any prompt engineering, since LLMs already listen past your typos and tangents. You don't start with kits, you start with directories: point Claude at a messy folder and ask it to build the beamer deck. Proof it conforms: typing /insights surfaced a workflow report claiming 1,642 hours of use. What transfers is only the idea that the mattress conforms; story, not documentation, unlocks adoption.
claude-codeai-adoptionendogenous-softwareresearch-workflowextensive-margin
TIER 4
May 13, 2026
Claude Code logs its full reasoning, including abandoned work, to a per-project JSONL "diary" — making agent mechanisms observable as text-as-data, unlike causal mechanisms, which are hard to identify when parallel-trends assumptions need not transfer across outcomes and unbounded heterogeneous treatment effects break OLS twoway FE, IV, and even Popperian falsification. In a minimum-wage study with primed agents, the diaries show specification searching: Claude runs a model, dislikes the estimate, and switches models. So humans must own verification and scaffold the target estimand; production advantage isn't verification advantage.
AI agentsspecification searchingJSONL logsheterogeneous treatment effectsverification
TIER 4
Mar 19, 2026
Completing cognitive tasks with an AI agent is not the same as learning; you cannot gain knowledge without resistance, just as you cannot build muscle without it. The Matrix "I know kung fu" fantasy that ChatGPT-4 seemed to promise is false — there is no free lunch on skill. AI works profoundly well in domains where you already have deep expertise, and jaggedly where you don't. Illustration: reviving an old paper using Callaway and Sant'Anna's difference-in-differences estimator on individual worker data. Claude ran a code audit and confidently insisted a "never treated" group was miscoded, demanding the whole pipeline be scrapped. Pushing back — forcing three-way verification against csdid as ground truth, stripping covariates, hand-computing ATT(g,t)s on tiny datasets — exposed Claude's actual error: it had been computing between-group differences (treated minus control means), a cross-sectional comparison, not a difference-in-differences at all. A parallel anecdote: a reasoning model wrongly told a workshop attendee double-robust estimation justifies different covariate sets. The danger is confident, professional-looking, runnable output that's wrong; human capital depreciates, so stay vigilant.
ai-and-expertiseverificationcode-auditcallaway-santannahuman-capital
TIER 4
Jun 10, 2026
Claude Code reproduces every habit that once kept an ADHD-inattentive researcher's error rate low—consistent naming, organized directories, automated tables, version control—yet introduces a new "drift": scripts never written down, analytical samples and treatment units silently shifting, errors surfacing outside their usual locations and uncatchable after a two-day gap. The conversational, co-authoring prompt style starves the scaffolding-memorization that compensated for forgetfulness. Fixes sought: 5-6 habitual policies, an analog notebook, and a dashboard "beautiful deck" reading progress logs.
AI for researchClaude Coderesearch driftverificationADHD/workflow
TIER 4
Jun 11, 2026
Because AI agents have "flattened the isoquant," cognitive output can now be produced with near-zero human time, but that path leads to a "depreciation trap": below a personal reservation time H-bar, you produce work you cannot verify, since human capital and the feeling of knowing depend on attentive time. Treat daily amnesia as the resting state and make zero errors a constraint, not a minimization goal that tolerates optimal nonzero error. The remedy is a harness dashboard built on four pillars: amnesia plus "love always remembers" (a 50 First Dates analogy — keep the project's discoveries within reach); narrative, because biography-of-econometrics stories jog memory (aided by aphantasia); beauty, which makes you stop and stare and thus stay informed; and a surgery/aviation-style 9-10 point checklist where nothing advances unless audited by /referee2 as replicable.
AI for researchClaude Coderesearch workflowhuman capitalverification
TIER 4
May 21, 2026
Stacking /skills atop an old folder-hierarchy workflow no longer fits once AI agents dominate research; the fix is a redesigned human-in-the-loop "harness" — the software infrastructure wrapping an LLM that manages context's lifecycle (intent, specification, execution, verification, persistence), everything except the model. Rebuilding starts from philosophy, not tools: rereading Allen's Getting Things Done. Open questions — where to record conjectures now that typing-flow is gone, how to visualize ideas, and the human/Claude property rights over tasks. Cites Shockley's multiplicative, log-normal model of scientific productivity.
AI agentsresearch harnessworkflow designClaude Codehuman-in-the-loop
AI agents and the crisis of academic publishing
3 tier-5 · 4 tier-4
If the marginal cost of a submission-quality manuscript falls toward zero, what happens to journals? This cluster is Cunningham's "fan fiction" of the publishing transition rendered as serious economics. Submissions surge on both margins, fixed acceptance slots drive accept rates toward 1%, the fixed referee pool can't scale, and the heuristics editors use to triage collapse because AI papers are both more numerous and (on average) better — the left tail of quality disappears. He works the formal machinery (supply and demand, Little's Law stock-flow identities, HHI of the publisher market) and then turns prescriptive: define the journal's objective function, LLM desk-screening, require runnable code repos at submission, raise Pigouvian submission fees with price discrimination. The threat to authors is concrete too — violate a dominant publisher's AI-disclosure policy and you may be banned from most of the market at once.
TIER 5
Mar 2, 2026
AI agents collapse the cost of producing a submission-quality academic paper to near zero, so the binding constraint on science shifts from production to evaluation. A full economics paper—shift-share identification, web-crawled data, analysis, write-up, referee revisions via refine.ink—now takes a couple hours and ~$100. With top-5 acceptance at 3-5%, the rational move is writing a hundred papers and submitting all; expected 5x submission volume (12,000 economists going from 3 to ~10 papers each). Reimers/Waldfogel found ChatGPT tripled Amazon titles, lowering average quality from the left tail while the frontier held; Zurich's Project APE has autogenerated 204 papers, winning 4.7% (rising to 7.6%) of head-to-head matchups against AER articles. The 3,800 fixed journal slots can't expand, so acceptance drops toward 1% or 0.5%; journals earn more in fees ($6.2M→$31M), referees (unpaid, ~54,000 capacity vs. 146,000 needed) break, desk-rejection rises to 90% via noisy pedigree heuristics. The equilibrium is a prisoner's-dilemma arms race—everyone spends $3,200/year to stay in place. Posting 75 unpublished manuscripts becomes a paper-mill signal that markets penalize.
ai-and-publishingsupply-and-demandpeer-reviewproject-apeevaluation-bottleneck
TIER 5
Mar 13, 2026
An accounting identity, not a model, dooms academic journals once AI cheapens manuscript production: by Little's Law, the stock of papers under review equals inflow times average review duration, so if AI doubles submissions while referees and editors stay fixed, queues must grow, wait times stretch, and manuscripts-per-referee must rise. The borrowed mechanism is Neal and Rick's prison "bathtub": longer sentences alone forced incarceration to triple (160 to over 500 per 100,000, 1970s-2008) by widening inflow while clogging the drain — the stock had to rise.
The baseline was already broken: Card and DellaVigna show top-5 submissions nearly doubled (3,000 to 6,000, 1990-2012) while published articles fell (~400 to 300), crashing acceptance from 15% to 6%; papers are now three times longer; 90% of econ PhDs never publish even half a top paper (Conley and Onder). AI worsens this on two margins — intensive (productive economists writing more) and the scarier extensive margin, activating the dormant 80% who currently produce nothing now that fixed setup costs vanish. A 2025 Science analysis found 36% of early-2024 submissions contained AI text, only 9% disclosed.
Desk rejection, the only fast lever, fails arithmetically: holding the queue at 3x volume means raising rejection from 50% to 83%, and the surviving papers are better-polished — clean code, every robustness check — so the signal that enabled fast triage (sloppy writing) disappears; editors would reject better science.
The prescription, addressed to editors: first decide the objective function (AI-prohibition-via-detection versus maximizing scientific innovation — pick the latter), then deploy LLM pre-desk screening, mandate runnable code repositories at submission (not just acceptance), and raise fees as a Pigouvian tax on referee time, with LMIC and early-career waivers. The long-run equilibrium adjusts; the transition (per Gans) is the danger.
peer-reviewstock-flow-identitylittles-laweditorial-policyai-and-publishing
TIER 5
Feb 20, 2026
Given a continuous -100 to +100 thermometer to score 285,376 congressional immigration speeches, gpt-4o-mini (zero-shot, temperature zero) spontaneously used only nine values — every score a multiple of 25 — collapsing a 201-point scale into a 9-point ordinal. This reproduces "heaping," the focal-point rounding humans show on survey feeling thermometers (95% of 2012 ANES responses rounded to multiples of 5), explained by Krosnick/Simon satisficing — yet the LLM has no cognitive load to satisfice. Boundary cases clustered near zero (reclassified means -25 and +2.5) versus -54/+48 for agreed cases, validating the earlier 69%-agreement, identical-aggregate-trend finding. Unlike Horton's "Homo Silicus" or persona-prompted work, nobody told it to act human; it absorbed measurement noise alongside content, illustrating Autor/Polanyi tacit knowledge. Four benchmark datasets showed accuracy (55-97%) tracks category separability, not tripartite structure.
LLM measurementtext classificationsatisficing / heapingtacit knowledgegpt-4o-mini
TIER 4
May 27, 2026
Journal AI-disclosure policies are demand-side market power, not neutral rules. Just as Match owns ~75 dating apps (HHI ~4500), violating Elsevier's or Wiley's AI policy locks you out of ~61% of economics journals—An, Williams & Xiao (2026) plus Claude yield an HHI of 2,430 (1,700 if Value in Health is dropped). Prestige association/university journals (AER, JPE) stay open, but not the modal economist's bread and butter. Disclosure that merely names AI use in methods is acceptable; a coding ban isn't. With detection rare, Becker's 1968 logic predicts harsh punishment. Coauthors must disclose too.
AI disclosure policyjournal market concentrationHHIBecker crime and punishmentresearch ethics
TIER 4
Mar 27, 2026
AI agents writing economics papers with zero human guidance reproduce the p-hacking signature of the human literature they trained on. The Social Catalyst Lab's APE project auto-generated 651 program-evaluation manuscripts (targeting 1,000) — real data, R scripts, estimators, figures, robustness tests. Cunningham used Claude Code to clone the repo and batch-submitted all 651 to GPT-4o for zero-shot classification (since the agents' own metadata was often wrong); the full run cost $12.28. GPT-4o found 61% diff-in-diff, with Callaway-Sant'Anna and TWFE the dominant DiD estimators, plus 87 "triple-diff" papers classified by the paper's rhetoric rather than its estimator. Agents spontaneously produce event-study plots, first-stage plots, density and balance tests, and rollout maps — never instructed to. 81% (528) explicitly name an estimand, ATT most common (422), with 45 naming LATE despite only 20 IV papers. The p-hacking is real: median t-statistic 1.94, a visible density spike at t=1.96, and 52% more mass just above the threshold than below (ratio 1.52, versus Brodeur et al. 2020's 1.4 for human top-journal papers). By method, IV bunches worst (3.5), then RDD (1.87) and DiD (1.5). With no publication incentive, the agents mimic a p-hacked corpus because that's what a paper looks like.
AI-generated papersresearch design classificationdiff-in-diffestimandsp-hacking (retracted)
TIER 4
Mar 5, 2026
When the marginal cost of a submission-quality econometrics paper collapses to zero, the binding question stops being who wrote it and becomes whether the finding is true. Overnight, from a vague prompt, Claude Code chose a topic (marijuana legalization's effect on employment and mortality), pulled real BLS/CDC WONDER and Harvard Dataverse data, ran Callaway-Sant'Anna diff-in-diff with robustness, and produced clean event studies — 3.5 hours active time, ~$100 of Refine.ink. Result: cannabis legalization appears to raise weekly wages ~2.2%, null employment, no overdose change. The right benchmark isn't AER (Zurich's Project APE: AI wins ~7% of head-to-heads) but field journals like JOLE, JHR, Journal of Health Economics — possibly within reach. Twenty years of causal-inference expertise caught a Sun-Abraham aggregation bug and bad incarceration data; the free-rider problem is that such verification skills will depreciate or never form. Is an unpublished event-study plot a fact? Should one even work on a machine-conceived paper?
automated-researchdiff-in-diffepistemics-of-factsfield-journalsverification
TIER 4
Jun 5, 2026
AI slop behaves like spit: your own chatbot conversations feel insightful, others' feel repulsive — the same material judged by source, mirroring the saliva experiment where people swallow their own spit but gag swallowing it from a cup ("cognition of disgust," an evolutionary disease-avoidance trait). If real, this AI bias makes others reject AI-packaged ideas regardless of merit. The fix is a blinded/revealed-authorship tournament (human-vs-AI, AI-vs-AI) measuring whether scores drop when AI authorship is exposed.
AI and repugnanceexperimental designAI writingbehavioral economicsresearch ideas
The Claude Code research harness — workflow, skills, and craft
1 tier-5 · 11 tier-4
The how-to spine of the AI series. Here Cunningham converts the verification thesis into concrete practice for empirical social scientists (not programmers): build external memory in markdown (CLAUDE.md, timestamped progress logs) to defeat the agent's amnesia, treat the agent as a thinking partner, verify via visualization, and run an adversarial "Referee 2" audit plus cross-language (R/Stata/Python) replication on the premise that hallucination is measurement error orthogonal across languages. The skill-building posts get specific — /split-pdf (chunk papers to cut hallucination), /beautiful_deck (decks as notes to your future self), /bibcheck, /blindspot, /tikz — and surface durable craft lessons: a "circuit breaker" to stop infinite compile-fix loops, the "marginal vs average user" risk of agents that act on your filesystem, and Deming's zero-error philosophy of codifying a rule so each defect can't recur.
TIER 5
Feb 2, 2026
Treat Claude Code as a thinking partner for empirical research, not a code-writing seal: the hard part is deciding what code to write and whether results mean what you think. Because Claude forgets everything between sessions, build external memory in markdown (CLAUDE.md, READMEs, session logs) it reads on startup. Use Socratic questioning ("guess what I'll ask next") to keep alignment, and verify via figures, not numbers alone. The core innovation is "Referee 2": open a fresh terminal, paste an adversarial-reviewer persona that runs five audits and files a formal report with major/minor concerns, never modifying author code. Confidence comes from cross-language replication—since hallucination resembles measurement error orthogonal across languages, matching R/Stata/Python results to 6+ decimals catches bugs single-language review misses. Slides are "sequential visual persuasion"; titles should assert, not label. Tools live in the public MixtapeTools repo.
claude-codereferee2cross-language-replicationresearch-workflowai-for-research
TIER 4
Jan 12, 2026
Claude Code is dangerous to empirical social scientists not because it goes rogue but because it acts rather than merely speaks: it operates your machine through shell commands, so anything you could do at the Terminal, assume it can and will do on your instruction. The economist's "marginal user" framing carries the argument. Today's average Claude Code user is a computer scientist who knows what `rm -rf` does; the marginal users now arriving in droves are quantitative social scientists whose mental model is graphical (RStudio, the Stata app), who may not know that R, Stata, git, and Python commands all run from the same shell, or that one command can wipe a directory, overwrite a drive (`dd`, `mkfs`), or destroy history (`git reset --hard`, `git push --force`, `git clean -fd`). The explainers don't cover this gap because they're written by average users for average users; the maps for social-science workflows don't exist because that work isn't sold on product markets, so no vendor draws them — researchers must map it themselves, fast.
The real risk is miscommunication, not malice. "Clean up this data and save it" might mean overwrite the original — violating the cardinal rule never to save over a dataset. "Move these files" may invoke `mv`, deleting them from source. Permission-fatigue from constant prompts makes you click yes into disaster.
The fix is a new "workflow" forming endogenously around the agent: version control everywhere, radical versioned backups plus originals kept on inaccessible external drives, habitual dry runs ("tell me what you'd run, don't execute") with sub-agent review, awkwardly narrow "annunciated" instructions, running first on a copied test environment, and always confirming your working directory before destructive operations. Take full ownership; blaming the AI is never available.
Claude CodeAI agent riskmarginal vs average usershell/Unixresearch workflow
TIER 4
Jan 14, 2026
Empirical social scientists, not just programmers, have enormous untapped gains from Claude Code — but it must be experienced to be understood. A 30-minute unscripted video walks through starting an old 2016 project on Texas House Bill 2's abortion-clinic closures (with undergraduate Andrea Schlosser). Four steps: a one-line title prompt triggers Claude to explore the folder, find all 105 data files plus Stata/R scripts and the manuscript, infer the timeline from timestamps, and summarize findings unprompted. Next, generate a README and a CLAUDE.md holding "rules of engagement" — never delete data or programs, stay within the directory tree, use a legacy folder, copy don't move — because crashed sessions lose all context unless written records exist. A self-created rules contradiction (moving into legacy) is resolved collaboratively by amending to a one-time move. Reorganization rebuilt 150+ files into a clean hierarchy; timestamped progress logs serve as workflow autosave. Claude Code is powerful enough to harm — "a rottweiler off its leash."
Claude CodeAI for researchresearch workflowCLAUDE.mdprogress logs
TIER 4
Jan 17, 2026
Beamer decks can replace note-taking: rather than slides for audiences, Cunningham has Claude Code build presentations that hand context to his future self and coauthors across work sessions. Because Claude trained on countless decks, it extracted the tacit "rhetoric of decks" into a deck.md — one idea per slide, titles as assertions, lead with conclusions, visual hierarchy. This illustrates Autor's Polanyi paradox ("we know more than we can tell") and Mollick's jagged frontier: LLMs are weak at precise calculation but strong at pattern extraction. The workflow: progress logs reconstruct context, rhetoric documents (deck.md, CLAUDE.md) encode preferences, beautiful outputs capture attention.
Claude CodeAI for researchtacit knowledgePolanyi paradoxworkflow
TIER 4
Feb 3, 2026
A skill is a reusable recipe: a `/command` triggers pre-written instructions (stored in `.claude/skills/<name>/SKILL.md`) so Claude executes a multi-step task without re-explanation. Make one only for workflows that are multi-step, repeatable, and fragile. Skills differ from personas like Referee 2, which must run in a separate fresh-context session to stay adversarial; a skill runs inside the current session. The `/split-pdf` skill solves two failures with academic PDFs: token-heavy documents trigger an unrecoverable "prompt too long" crash that wipes session context, and long reads degrade attention so Claude hallucinates results. It downloads the paper (never deleting the original), splits it into 3-4 page chunks via PyPDF2, and reads three splits at a time, extracting eight dimensions (research question, audience, method, data, statistical methods, findings, contributions, replication feasibility) into a running `notes.md`. Shorter, repeated engagements decorrelate hallucination errors. Demonstrated on Gentzkow, Shapiro & Sinkinson (AER 2014).
claude-codeskillssplit-pdfhallucinationpaper-reading
TIER 4
Jan 29, 2026
Claude Code builds lecture slides not by knowing LaTeX but by having absorbed the tacit knowledge of skilled deck communicators that no one writes down. The method is "dictation," not vibe coding: talk through pedagogy, audience, big-picture outline and minutiae, then continuously tweak, rearrange, and scrap as each slide appears. Vaguely-held ideas work too — a request that each slide hold the same "marginal-benefit-to-marginal-cost ratio" for "optimal rhetoric" gets understood and attempted. Mid-deck, Claude wrote its own Beamer .sty theme from scratch and produced ambitious TikZ graphics (a filing cabinet drawing) never explicitly requested. Framed as supply-and-demand: Claude Code shifts both deck-production curves — marginal benefit up, marginal cost down and flatter — so time-adjusted slide quality rises. It is comparative advantage, narrowing the gap to naturally gifted course designers like Rebecca Thornton, not surpassing them. No prompt-engineering skill remains to learn; the tool serves anyone whose work lives in a directory of folders and files.
claude-codedecksteachingrhetoric-of-decksproductivity
Claude Code Part 13: I Asked Claude to Replicate a PNAS Paper Using OpenAI's Batch API (Part 1)
TIER 4
Feb 5, 2026
The setup half of the PNAS replication: Claude Code web-crawls the replication package, builds a self-contained project structure, designs the classification prompt, chunks 305k speeches into JSONL batch files, estimates cost at $11, and runs a Referee 2 audit that catches label-normalization edge cases and missing Cohen's Kappa before submission. A useful end-to-end walkthrough of orchestrating a hard empirical task with an AI agent, including defensive scripting and pre-run code review.
claude-codereplicationbatch-apireferee2research-workflow
TIER 4
Feb 23, 2026
Every AI error is information: don't just fix the defect, find why it happened and change the workflow so it can't recur — Deming's postwar lesson applied to individual knowledge work. Running `/insights` over 73 Claude Code sessions (585 messages, 44,486 lines written) produced a "portrait" naming the author's edge as "ambitious delegation with sharp correction" — delegate heavily, then audit aggressively; the 82% success rate came because of corrections, not despite them. Building a Beamer deck surfaced TikZ errors LLMs can't see because they lack eyeballs and are bad at spatial reasoning (like chess after fifteen moves). The fix converts spatial problems to arithmetic: a Bezier curve's depth is `(chord/2) × tan(bend_angle/2)`, so Claude computes rather than eyeballs. Each failure category (Bezier curves, crossing arrows, overlapping rectangles) became a rule; `tikz_rules.md` grew to nine rules and a five-pass workflow — "prosthetic spatial reasoning." Converting `/compiledeck` from a command (a memo Claude reads once) to a skill (a structured directory training how) entrenches zero-tolerance. Solutions are personal; don't download others' starter packs.
AI workflowDeming / zero errorskills vs commandsspatial reasoning/insights
TIER 4
Apr 13, 2026
A skill that audits output after generation can't fix problems baked in at generation. The `/beautiful_deck` skill produced gorgeous slides but high TikZ error rates because it told Claude what to check, never how to generate safely; the fix adds six generation rules (explicit node dimensions, directional edge labels, no `scale` on complex figures, parameterized styles in the preamble not Beamer frames) plus a "circuit breaker" halting after three failed fix attempts instead of spiraling for an hour. `/split-pdf` gained reader-contributed agent isolation (PDF rendering in subagents to dodge context-size limits) and persistent `_text.md` extraction for reuse. `/blindspot` (formerly `/fletcher`), drawing on Shklovsky's "make the stone stony again," uses a 2x2 vice/virtue grid to catch what you stop noticing; it runs in-session before `/referee2`, which audits implementation in a fresh session.
Claude CodeskillsTikZ debuggingresearch workflowagent isolation
TIER 4
May 6, 2026
Make your own skills rather than borrow them: copying others' CLI snippets is a leaky pipeline that will eventually smuggle malware, whereas letting Claude work from URLs stays safe, and skills are functions of your own human capital—not Kung Fu downloaded like Neo. The organizing premise is "gradient decay" (diminishing returns as token/task load grows): chunk work into tiny single-purpose agents to dodge it. /split-pdf cuts an N-page PDF into N/4 four-page pieces, spawns one agent per piece to summarize, then a final agent merges—avoiding the choke a 100-page PDF causes. /bibcheck audits citations after the Sullivan & Cromwell hallucinated-citation scandal: Case 1 spawns one agent per citation, Case 2 one agent per bibfield (title, year, journal). Agents write referee reports, never auto-correct. Earlier /tikz once looped hundreds of times and maxed out tokens; he stripped it back. A /split-pdf accuracy experiment is still pending.
AI skillsmulti-agent designgradient decaysplit-pdfbibcheck
TIER 4
May 7, 2026
The AI demo that lands on skeptical research audiences is not "watch Claude write a paper" — it's having Claude build a presentation live, with the human performing it. At the Harvard Kennedy School, an empty "Kennedy" folder and one typo-ridden live prompt (run with --dangerously-skip-permissions) drove sub-agents to split Kremer and Levy's 2008 dorm-roommate peer-effects paper, summarize it, simulate its regression tables in R, and compile a 20-slide Beamer deck in 15-30 minutes while the talk proceeded. Decks dodge the "Luddite" repugnance manuscripts trigger: they're already shared, instantly verifiable from one's seat, and reconstructing published coefficients as faithful simulated figures is technically harder yet easier to receive. The other lesson: nothing is lost — every session lives in ~/.claude/projects/<dir>/<session-id>.jsonl (the Kennedy one ran 445 messages, 3.7MB), a fully auditable, reproducible "flight recorder for thought." The AI is the medium; the human is the rhetor.
AI demosbeautiful deckssession logsreproducibilityrhetoric
TIER 4
Apr 24, 2026
Use an AI agent for a task when four conditions hold together: it is high-value, time-consuming, hard to do well even with infinite time, and easy to do badly or wrong. Referee reports satisfy all four — they are the backbone of peer review, devour attention, demand forgetting the author's identity, and are trivially botched by idiosyncratic bias. The agent does not write the report; it builds an executive map so the human reads the paper faster. The workflow chains custom skills. /split-pdf chunks the PDF (~4 pages) into machine-readable markdown — PDFs are "hieroglyphics," not text — extracting research question, target parameter, identification assumptions, and core evidence. /beautiful_deck builds a Beamer deck for an audience of one, leaning on narrative and on simulations that mimic the paper's estimator and data; such a simulation exposes, e.g., that two-way fixed effects is maximally biased when a federal minimum-wage hike removes all untreated controls, and that Callaway–Sant'Anna cannot even run there. /referee2 critiques the agent's interpretation (not the manuscript), /blindspot hunts non-headline errors like sample sizes that don't add up, and /tikz checks label collisions via Bézier math (~50% success). Verification, not production, is now science's bottleneck.
AI agentsreferee reportsClaude Coderesearch workflowverification
Continuous-treatment diff-in-diff and the TWFE decomposition
1 tier-5 · 6 tier-4
A self-contained tutorial series in which Cunningham teaches himself (and the reader) the Callaway–Goodman-Bacon–Sant'Anna continuous-treatment estimator by building it from the ground up with Claude Code. The throughline is one durable conceptual point: a single TWFE/FWL coefficient under a continuous dose can be algebraically rewritten as several different weighted averages — levels, scaled levels, causal response, scaled 2x2 — each answering a distinct question, none cleanly equal to the causal estimand, and negative weights are the price of clean untreated-vs-treated comparisons. The arc moves from the Frisch-Waugh-Lovell derivation through the four-piece levels decomposition to an interactive R Shiny app that visualizes the sign-flip where below-mean-dose units get negative weight. The motivating slogan: the regression never changes; the question does.
TIER 5
Apr 23, 2026
Table 1 of the Callaway–Goodman-Bacon–Sant'Anna continuous-treatment DiD paper decomposes only one TWFE regression, not four; its four rows are algebraic rewrites of the same β-hat, each expressing that single number as a weighted average of a different underlying parameter. Using Lu and Yu's (2015) China-WTO-tariff regression, the rows answer four distinct questions: the level effect (treated-at-dose vs untreated), the marginal slope (derivative, CBS's ACRT), the per-unit scaled effect, and pairwise dose-to-dose slopes. None cleanly equals the causal estimand a researcher writes down upfront. The level ("clean") decomposition forces weights to sum to zero, so below-average-tariff industries get negative weights via FWL recentering. The scaled 2×2 row avoids negatives but compares two treated units—Bacon's "forbidden comparison." Negative weights are the price of clean untreated-versus-treated contrasts, not a bug. OLS cannot tell you which question you asked. The paper's constructive move: reason population-first, define the target parameter, then forward-engineer an estimator under explicit identification assumptions.
continuous diff-in-diffTWFEFrisch-Waugh-Lovelldecomposition weightsestimands
TIER 4
Apr 9, 2026
Applied researchers won't adopt a new diff-in-diff estimator until shown their current one is broken — Goodman-Bacon's 2021 result that two-way fixed effects (TWFE) is biased even under parallel trends is what made differential-timing estimators stick, and the same wedge motivates learning the continuous-treatment paper by Callaway, Goodman-Bacon and Sant'Anna ("CBS," conditionally accepted at AER). The strategy is "backwards engineering" (per Sant'Anna): run the regression, then crack the coefficient open via Frisch-Waugh-Lovell to see what estimand it actually targets — chosen over "forwards engineering" because people must first learn what the coefficient they love means. The application is Lu and Yu (2015, AEJ:Applied) on China's 2001 WTO accession, using predicted tariff cuts as a continuous dose to show larger cuts reduced within-industry markup dispersion. The build uses Claude Code skills: /split-pdf (chunk papers, summarize each, then summarize the whole), /beautiful_deck (Beamer slides on Aristotle's ethos/pathos/logos, one idea per slide), /tikz (fix label/arrow collisions via Bezier depth formulas), and /referee2 (fresh-terminal adversarial deck audit). Part 1 only builds architecture and the deck.
continuous diff-in-diffTWFEBacon decompositionClaude Coderesearch workflow
TIER 4
Apr 15, 2026
A two-period TWFE regression of outcomes on unit/time fixed effects and a continuous dose (e.g. how much a municipality raises the minimum wage, not just whether) reduces, via Frisch-Waugh-Lovell, to the OLS slope of the unit-level first difference on dose. Working the "Levels" row of Callaway, Goodman-Bacon and Sant'Anna's Table 1: condition on D, split its mass at zero from its positive support, add and subtract m(0). The m(0) terms cancel because mean deviations sum to zero, leaving the coefficient as an integration-weighted average of dose-level calculations—still purely algebraic, no causality yet.
continuous diff-in-diffTWFEFrisch-Waugh-Lovellderivationdecomposition weights
TIER 4
Apr 20, 2026
Callaway, Goodman-Bacon and Sant'Anna's continuous-dose diff-in-diff paper decomposes the TWFE coefficient via Frisch-Waugh-Lovell; the "level" weight has three ingredients — mean dose E[D]=0.164, variance 0.0202 (denominator), and density f_D(l), integrated over doses. The decisive feature is a sign flip: the weight is zero exactly at the mean, positive for above-average doses, negative below. A Claude Code-built shiny app sliders the dose to show this; a Gaussian kernel spuriously smeared density left of the smallest observed dose, fixed by adding min/max dashed lines.
continuous diff-in-diffTWFER Shinydecomposition weightsClaude Code
TIER 4
Jan 20, 2026
Reviving a 2019 Journal of Human Resources paper (Cunningham, Schlosser, Lindo, Myers) on Texas HB2 abortion-clinic closures, using Claude Code to re-estimate distance effects under Callaway, Goodman-Bacon, and Sant'Anna's conditionally-accepted AER continuous diff-in-diff method. Claude audited two rival distance datasets: an earlier thesis version backdated 2010 distances to 2006 (assuming no pre-period closures, killing within-county variation) versus Caitlin Myers's hand-tracked "ground truth." They diverge 8% of the time by missing New Mexico/Oklahoma clinics—Lubbock reads 307 miles versus 78. Beamer/TikZ decks, CLAUDE.md, todo.md, and logs serve as memory. Unsolved: urban counties (42% of Texas) supply no treatment variation, breaking the counterfactual.
claude-codecontinuous-diff-in-diffabortion-accessidentificationdata-audit
Claude Code Series (part 8): Resurrecting and Extending an Old Abortion Paper Towards Using Continuous Diff-in-Diff [original send]
TIER 4
Jan 20, 2026
Identical content to 0111: a Claude Code case study reviving the JHR Texas HB2 abortion project for continuous diff-in-diff, with a thesis-vs-JHR distance data audit and the urban-counterfactual identification problem. This is the original post that was later deleted and reposted as 0111 because it had the wrong video attached; substantively it is the same essay.
claude-codecontinuous-diff-in-diffabortion-accessidentificationdata-audit
TIER 5
Jun 12, 2026
Parallel-trends is just selection bias on a derivative: where simple comparisons carry selection bias as the gap in mean untreated potential outcome Y(0) between treatment and comparison groups, diff-in-diff's "non-parallel trends bias" is the same gap in first differences of E[Y(0)]. Working the 2x2 in four moves (expectations, swap realized for potential outcomes under no-anticipation, add a zero, rearrange) isolates the ATT plus that lone bias term — proof the calculation isn't diff-in-diff without the assumption. The bias term is itself a 2x2 on Y(0). Following Imbens, the horizontal (first-differences) and vertical (within-time group-difference) regressions are algebraically identical; TWFE does both. Cites the JEL Practitioner's Guide and Mostly Harmless Econometrics.
difference-in-differencesparallel trendsselection biassynthetic controlvertical regression
Callaway–Sant'Anna and the perils of staggered DiD with covariates
3 tier-5 · 4 tier-4
A sharp, practitioner-facing cluster on a single underappreciated fact: Callaway–Sant'Anna with covariates secretly fits one propensity-score logit *per treatment cohort*, and that hidden stage is where things break. Cunningham develops the consequences across several posts — Peduzzi's events-per-variable rule means the binding constraint is treated units per cohort (so U.S. state-level staggered panels with singleton treated states routinely violate it), some packages quietly drop covariates or fail to converge, and an apple-to-apple audit running identical specs across six R/Stata/Python packages produced ATT estimates ranging up to ~5x apart, driven by covariate handling, matrix conditioning, and near-separation in the logit. The constructive payoff: zero-covariate baselines, z-scoring covariates, reporting package and version, and switching to regression adjustment (which has no events-per-variable problem).
TIER 5
Apr 1, 2026
The Callaway and Sant'Anna staggered-adoption diff-in-diff estimator (11,000+ cites, the default AI agents pick) cannot absorb as many covariates as users assume, because its IPW version estimates one logit per treatment cohort, not one regression overall. Covariates enter through propensity scores fit separately for each of the (say) 14 treatment years; what constrains each fit is treated units in that cohort, not the total. Peduzzi et al. (1996)'s "ten events per variable" rule then bites locally: a 2003 cohort with 47 treated units chokes on covariates a 2006 cohort with 216 handles fine, risking flat likelihoods, overfitting, and perfect separation. Worse, some packages silently drop all covariates and regress treatment on a constant. Run the 14 logits yourself; state panels with 50 units almost always trigger this. Regression adjustment sidesteps logits but trades one problem for another.
Callaway-Sant'Annadiff-in-diffpropensity scorecovariatesevents per variable
TIER 5
Mar 12, 2026
Six independent software packages implementing the identical Callaway and Sant'Anna (2021) estimator on identical data and specifications return wildly different answers — ATTs from 0.00 to 2.38 homicides per capita on a mental-health-closure question, one package saying half a standard deviation, another more than double. Claude Code wrote ~96 scripts (16 covariate specifications across did, ddml, csdid, csdid2, differences, diff-diff in R/Stata/Python) on Brazilian municipal data, 2002-2016. With zero covariates all packages agreed to four decimals (ATT ≈ 0.31); adding one variable, poptotaltrend (population × year, reaching ~10 billion), fanned estimates out. A two-way ANOVA: 40% of variation from specification, 16% from package, 44% from interaction. Two culprits: floating-point matrix-inversion failure from huge condition numbers (fixable by z-scoring), and near-separation in the propensity score — outlier structure scaling doesn't fix, since each package's logit optimizer (MLE vs. csdid's inverse-probability tilting) handles perfect prediction differently. R's did silently returned zeros. This is undocumented publication bias: package choice is a substantive decision no journal requires reporting. Recommendations: report package/version, standardize covariates, run a zero-covariate baseline, and use AI agents for such cross-language code audits.
callaway-santannapackage-variationpropensity-scorenear-separationcode-audit
TIER 5
Jun 8, 2026
Automating applied econometrics raises, not lowers, the returns to econometrics knowledge, because automation replaces production but not verification. Letting Claude "rip" with /skills backfires: it specification-searches, silently writes up whichever method fits the prompt, and buries decisions in a JSON. A four-hour Callaway-Sant'Anna (CSDID) debugging ordeal—NAs in 2x2s—turned out to be the CRAN `did` build versus the newer GitHub one; two Claudes confidently ran down wrong conjectures throughout. Agents make reasoning mistakes, unlike human coding mistakes, and strip away the "epistemological feeling" of knowing what produced what. The fix is structured analog workflows—checklists (Gawande's Checklist Manifesto; "Pedro's Checklist"): name the estimand in potential outcomes, tabulate treated units per cohort (7-10 events per covariate for the propensity score), pre-register aggregation, plot rollouts. Constraint: zero error.
AI for researchCallaway-Sant'Annastaggered diff-in-diffchecklist/workflowestimand specification
Claude Code 34: Using "Dispatch" on my phone with Claude Code to revisit the cannabis paper
TIER 4
Mar 20, 2026
Framed around the new phone-based "Dispatch" remote-control feature, but the substantive core is a methods lesson: Callaway-Sant'Anna with covariates secretly fits a per-cohort logit propensity score, and Peduzzi et al.'s "events per variable" (EPV) rule means you need ~10 treated units per covariate per cohort-year to avoid biased propensity scores. Cunningham argues U.S. state-level staggered DiD routinely violates this (singleton treated states), and the fixes are dropping to county-level data or switching from doubly-robust/IPW to regression adjustment, which has no events-per-variable problem. Useful for anyone running CS with covariates.
callaway-santannaevents-per-variablepropensity-scoreconditional-parallel-trendsclaude-code
TIER 4
Feb 25, 2026
Treat LLM coding hallucinations as classical measurement error: random, language-specific syntax slips that systematically corrupt downstream results. A real example—Stata's `replace olddog = 10 if olddog>10` silently recodes missing values too unless `& olddog~=.` is appended—shows how one bad random draw cascades through a pipeline unnoticed since the code still runs. The fix exploits independence: if errors across R, Python, and Stata are uncorrelated (Cov=0), the chance all three hallucinate identically is the product of tiny probabilities. So have Claude Code aggressively audit code like a health inspector AND fully replicate the pipeline in two other languages, demanding identical tables to several digits. This works for deterministic methods (OLS, diff-in-diff, IV) but not random ones (bootstrap, MCMC, ML). Case study: five diff-in-diff packages (csdid, csdid2, did, differences, diff-diff) on a Brazilian CAPS deinstitutionalization study.
code auditdifference-in-differencesmeasurement errorLLM hallucinationreplication
TIER 4
Feb 27, 2026
Hand the same DiD problem to fifteen isolated AI agents and the spread of estimates becomes a cheap robustness audit — an automated version of the "multi-analyst" findings (Silberzahn's 29 teams; Huntington-Klein's seven economists, whose cross-analyst SD ran 3-4x the typical standard error; Menkveld's "non-standard errors"; Borjas-Breznau, where ideology predicted the effect's sign). The setup: 15 fresh `claude -p` agents, three each across five Callaway-Sant'Anna packages (Python's differences/diff-diff, R's did, Stata's csdid/csdid2), all run on Dias-Fontes's Brazilian CAPS mental-health rollout over 5,476 municipalities, under strict isolation. Result: every structural choice was unanimous (not-yet-treated controls, universal base period, no trimming, 15/15). Variation lived entirely in covariate selection — log GDP near-universal, geographics rejected, but the confounder-versus-mediator line (poverty 10/15, health spending 7/15, Bolsa Familia 2/15) shifted agent to agent. Whether that disagreement actually moves the ATT is held for Part 2.
difference-in-differencesCallaway-Sant'Annanon-standard errorsmulti-analyst designAI agents
TIER 4
Mar 4, 2026
Covariate selection, not the choice of estimator, drives most variation in Callaway-Sant'Anna difference-in-differences estimates. Across 15 runs on one fixed dataset through five language-packages (Python's differences and diff-diff, R's did, Stata's csdid and csdid2), all positive, the average ATT sat near 0.4, but an R agent using only state fixed effects got 0.17 while a diff-diff agent with 8 covariates got 1.83 — nearly 11 times larger, each rationale defensible. Roughly 77% of variation came from between-package differences. The "non-standard error" punchline: the standard deviation across the 15 point estimates (0.442) is 2.4 times the average reported standard error (0.185) — sampling-based standard errors capture none of the team/discretion uncertainty. Picking g-5 over g-1 as baseline silently changes the parallel-trends assumption (always estimated as long differences). Next: doubleml added, 20 agents per package for 120 estimates, free covariate range.
callaway-santannamany-analyst-designcovariate-selectionnon-standard-errorscode-audit
Applied empirical essays and the sociology of the discipline
0 tier-5 · 5 tier-4
The remaining substantive essays that sit outside the methods tutorials and the AI-workflow spine: applied economics, the statistics of research integrity, and the intellectual history of the field. Cunningham dispatches Claude Code to crawl EDGAR and build an HHI analysis of Match Group's dating-app portfolio (a zero-price market that evades antitrust); works through why reconstructing t-statistics from rounded coefficients manufactures false bunching at t=2 (and publicly retracts his own p-hacking claim when he finds the artifact); distinguishes three sources of uncertainty — sampling, design-based, and analyst/researcher uncertainty that standard errors never capture, now operationalizable via many-analyst automation; and maps the family tree of causal inference through Orley Ashenfelter's academic lineage.
TIER 4
Dec 31, 2025
Modern causal inference in economics descends from Orley Ashenfelter and Princeton's Industrial Relations Section, a quantitative-labor lineage that merged with Don Rubin's potential-outcomes framework once Imbens and Angrist linked instrumental variables to the Rubin model. Three streams—Princeton's natural experiments, Harvard's experimental design, Chicago's Heckman/structural—form "two rivers converging." A genealogy database (~1,100 economists across four generations: Card, Heckman, Angrist, Duflo, Abadie) crowdsources missing advisor-student edges, source-verified, treating people as the real confounders.
causal inference historyOrley Ashenfelteracademic genealogyPrinceton IRScredibility revolution
TIER 4
Apr 6, 2026
Match Group owns Tinder, Hinge, OkCupid, Plenty of Fish and more, controlling roughly 65% of a market where over half of couples now meet online; a ban from one app ejects you from all. Revenue-share HHI runs 4,600–5,600 versus the 2,500 "highly concentrated" threshold, yet DOJ ignores it because antitrust's price-focused consumer-welfare framework can't see zero-priced products. Match acquired 25+ firms with essentially no merger review—a regulatory failure if dating is critical social infrastructure.
antitrustmarket concentrationzero-price marketsonline datingClaude Code research
TIER 4
Apr 2, 2026
Extracting coefficients and standard errors from published tables to reconstruct t-statistics produces false p-hacking evidence, because the rounding done purely for display collapses the continuous t-stat ratio into discrete heaps near 2. Authors' actual hypothesis tests use unrounded numbers, so their p-values are unaffected; only third-party reconstruction inherits the artifact. Ratios of exactly 2 are the second-most-common single-digit-integer ratio (after 1), so small-scale outcomes—coefficients like 0.01527 with leading zeros—heap there once rounded; large-scale outcomes (income on a college dummy) don't, since rounding only touches fractional digits and t-statistics are scale-invariant. Brodeur et al.'s diff-in-differences p-hacking signal vanished after a corrected fix (IV survived); the same flaw afflicted Cunningham's DiD-heavy APE-project analysis.
p-hackingrounding artifactst-statisticsShiny appresearch methods
TIER 4
Mar 30, 2026
A t-statistic spike near 1.96 in 651 AI-written economics papers wasn't p-hacking; it was an artifact of dividing rounded inputs. Because papers report coefficients and standard errors rounded (e.g., 0.035/0.021 = 1.67, but 0.04/0.02 = 2.0), reconstructing t-stats from those LaTeX-table values collapses many "near-2" cases onto exactly 2.0. Small numbers dominate in labor economics (log outcomes, proportions, linear-probability coefficients), and 2:1 is the simplest small-integer ratio, so the heap lands at 2, not 1.96. A simulation of smooth, unmanipulated data still piled 7.2% of t-stats onto 2.0. Brodeur et al. (2020) avoided this by using software-output t-stats, not reconstructions. A donut-hole drop of exact-2 cases flattens the reported Brodeur ratio from 1.52 to 1.02—no bunching.
p-hackingrounding artifactscorrectionBrodeur testAI-generated papers
TIER 4
Mar 6, 2026
A third source of statistical uncertainty exists that standard errors never capture: variation from who the researcher is. Sampling inference holds the researcher fixed and resamples the population; design-based inference (Fisher's lady tasting tea, randomization inference) holds the sample fixed and permutes treatment assignment. Both ignore that the analyst is not a transparent pipe. Silberzahn's many-analyst study sent one dataset and question to 29 teams and got a spread of estimates from defensible but subjective choices alone. A pipeline with ten decision points, three reasonable options each, yields 3^10 ≈ 59,049 possible estimates, none normally reported. Claude Code makes this operational: running the same Callaway–Sant'Anna diff-in-diff across five packages and varying covariates showed 75% of total variation traced to whether Python, Stata, or R was used. The hard unsolved part is automatically enumerating the discretionary (endogenous) nodes where reasonable analysts diverge, then perturbing them to build intervals and a p-value—how often is a node pivotal?
non-standard-errorsmany-analyst-designresearcher-uncertaintyinferenceclaude-code