Purity Tests for Academic Research Are Dirty
Disquieting trends on LLM use in research I hope stay marginal
Thanks for reading or listening! You can support Foreign Figures by liking, sharing, buying me a coffee, or becoming a free or paid subscriber.
When it rains, it pours. Recently when I opened my podcast app to listen to the morning headlines from NPR, I faced an onslaught of news relevant to this newsletter: new tariffs, a new escalation in the US-Iran War (which now is back on pause again), a possible nuclear agreement between the US and Saudi Arabia, and yet more volatility in global oil prices. And I shouldn’t forget the record-setting forest fires in Europe and the ones in Canada currently making trouble for those of us south of the border. Spoiled for choice, I’ve decided to talk about something else entirely.
I’ve already said my piece a few times on tariffs, and I’ve said a thing or two about the US-Iran War, so I’ll give these a pass. My theses on both remains the same. Tariffs are ineffective and counter-productive, and however long the US-Iran War drags on, the US isn’t likely to have a clear victory.
I haven’t talked about nuclear weapons or oil prices. Maybe I will in the coming weeks. But I need more time to think about nukes and the US-Saudi deal, which isn’t even signed as of my writing. It’d be premature to spend too much time on it now. As for oil, I would recommend checking another data-driven Substack by Robin Brooks all about macro economic trends (including oil prices). He’s said far more of value on the subject than I could ever hope to.
As for forest fires, I plan to talk about them next week, and I already have some visuals put together. So stay tuned.
Now that I’ve eliminated the immediately obvious choices, that leaves something else that’s been on my mind. It’s a nagging irritation that I’ve been struggling to put into clear words, one that’s likely overblown, but I think now is the time to try articulating it.
Okay, deep breath….
…I’m going to write about “AI” in research.
I’ve said before that have disdain for the term “artificial intelligence” because it is both vague and misleading. It doesn’t describe what the tool does, and it invites anthropomorphism. So instead I’ll use the LLM acronym, which stands for large language model. LLM also isn’t perfect. My first choice is to use the term “pattern engine,” but since I have no influence over names we as a society give to things, I’ll stick to LLM because most people know what I’m referring to, while it lets me avoid using “AI.”
Anyway, I have no plans (today) to complain about LLMs, much as the previous paragraph and my past takes might prime you to think. Rather, I want to share my thoughts about a trend I see in the academy surrounding research and these tools. It’s actually two trends running directly at odds, despite sharing roots in a common pathology.
I’m talking about anti- and pro-LLM purity tests, particularly of the agentic variety.
In one corner, you have the anti-LLM crowd. Call them “luddites,” or whatever you will. They hold the position that any LLM use, most seriously in the case of writing prose, but also with respect to writing code, is terrible and should be minimized. In the other corner, you have the pro-LLM crowd. Call them “sell-outs,” or some other name. They hold the position, not just that anyone can use agentic LLM tools in their research, but that everyone should.
Note that the “pro-” and “anti-” prefixes are not, at least to my mind, synonymous with “use” or “avoidance,” respectively (though they are probably sufficient conditions). I’m talking about an attitude toward these tools so strong, it compels puritanical evangelism. Plenty of folks can, and do, use these tools, or don’t, and can live and let live. My beef lies with those who can’t live and let live.
Some of what started this reaction in me was the varied responses I witnessed to research published by Anthropic, which sells one of the more ubiquitous agentic LLM tools, Claude Code. Anthropic published a report showing that only 20% of social scientists had adopted LLM agents in their work. I pasted a data visualization from the report below. Economists, and my fellow political scientists, are ahead of most other fields on LLM adoption, and agentic LLM adoption in particular. A quarter of political scientists surveyed said they use agentic tools, and 85% said they use LLMs in general.

Pro- and anti-LLMers had things to say about this research. Some rightly complained that this wasn’t a representative survey, so the results are biased and likely in favor of LLM users. One of the more controversial takes came from those in the pro-LLM camp shaming social scientists for their reluctance to adopt agentic tools in larger numbers.
When I see these results, I just see researchers making different decisions. Some want to adopt agentic tools. Others don’t. Some refuse to use LLMs at all, while some like to get their feet wet from time to time. All cool with me.
But the vocal extremes seem like they’ll only be content once everyone else does as they do, and their smugness is showing. Their attitude takes me back to my college days when my own smugness shown through, in my case, in the realm of health and fitness. Join me on a trip down memory lane.
When I was in college, I had my own dalliance with puritanical evangelism. Mine was of the health and fitness variety.
I had been overweight for most of my childhood and teenage years. That all changed my freshman year of college. I did a low carb diet, started lifting weights, and running. I had lifted weights for high school football since I was 15, but I didn’t take it terribly seriously and didn’t know what I was doing. Now I took lifting (and dieting) seriously. Perhaps too seriously.
After seeing good results, I feared losing progress. I went off the deep end. I got into the paleo diet and intermittent fasting, I obsessed over the design of weight lifting routines and the right calorie intake and macronutrient ratios, and I feared letting any unclean food touch my lips. When others in my dorm were playing Mario Cart in the lounge and eating ramen, I was in my room fasting, poring over fitness blogs, following links to randomized trials and population studies, and writing my own blog doling out my newfound wisdom (which today I now hope has been long lost to the ether).
I was, of course, a vocal proselyte of my new lifestyle to anyone who would listen. Why wouldn’t I be? At 5’9” I had gone from almost 230 to 170 lbs, could do weighted chin-ups, felt physically better than I ever had before, and gained some newfound confidence because of my changed physique. Everyone, EVERYONE ought to do what I was doing!
I was too young and foolish to realize that I had only found one path to health and fitness. I could have gotten similar results taking any number of other approaches, without obsessing over purity, or walking around with all-too-much arrogance in my head or fear in my heart that one minor slip-up would reverse all my progress in an instant. I needlessly deprived myself of the fun, food, and carefree attitude that defines undergraduate living, and more than a decade later I regret it.
I now like to think I have a better relationship with food and my health. Eventually, I started listening to more reasonable voices in the health and wellness space who taught me about the folly of focusing too much on identifying the right process for getting healthy and fit, and of focusing too little on what my health and fitness goals ought to be.
I give myself some grace, though. I had a profound and positive transformative experience, but I learned the wrong lessons from it. I fell prey to the dual lies that I had found something new, which was false, and that sameness with my approach was a necessary condition for promoting health, which was also false.
Too often things I hear staunchly pro- or anti-LLM academics say take me back to these wayward college days—or, should I say, what I see them post, because only by being online do I actually encounter those in either camp in large enough numbers to trick me into thinking I should care about this issue at all. This insight probably matters more than anything else I have to say in this post.
As with me and my fitness journey, many have had a profound experience, positive or negative, when encountering and using these tools. Regardless of the vector of their experience, an extreme but vocal few have falsely assumed their discoveries are new, and misguidedly decided sameness with their own process is the benchmark for good research practice.
Do those in either camp, when pressed, disagree about whether academics should produce quality, replicable, rigorous work, that they are willing to sign their names to? I think not. But instead of focusing on these metrics, the most extreme luddites and sell-outs choose to obsess over process. It seems the horseshoe theory doesn’t just apply to politics.
One example of the kind of experience that would have an evangelism-inducing-effect similar to my own health transformation is the sudden ability agentic LLM users possessed to organize their files and research project directories in a sensible way using a prompt rather than having to endure the drudgery of doing the organization by hand. In a previous post I called this the Marie Kondo Effect.
Organization is a great outcome from using these tools, but you don’t need them to get it, and the alternatives don’t all have to be labor intensive. I wrote my own script for auto-populating new projects with boilerplate files, organized into folders with clear names, and it has done wonders for my sanity when I do my work. Similar solutions exist and did before Claude Code was a twinkle in Dario Amodei’s eye. My solution, and others, require more proactivity than an agentic tool, but they don’t require paying for an agentic LLM subscription. Whatever approach you choose, the goal should be the same: an organized file structure that makes it easy for others, and your future self, to replicate your research.
On the other side, there are those who encounter these tools and fear the loss of the human in the research process. There is something innately fulfilling about doing research, or so I think. But there are lots of aspects of doing research short of using agentic tools that are only different in degree, not kind. When I do an analysis for this newsletter or my own academic research, I’m not cleaning data with a scrub brush, producing data visualizations on paper with colored pencils, or estimating regression models with a loom. I write code. And that code is just a set of instructions I give to my computer to tell it to clean data, make visuals, and estimate regression models for me. Prompting a coding agent to do the same task isn’t that much different. The basic work is the same: you sit at your computer and issue commands, stuff that you may not fully comprehend happens before your eyes, and then you evaluate the output.
I think some could do with a healthy dose of perspective. To dwell on the point a bit longer, if you think writing code yourself is toil that all quantitative social scientists would be fools not to escape, let me tell you about my experience working for nine summers for a commercial landscaper in hot humid summers, between two-a-day football practices in high school, sometimes working 12 hour days, once or twice on the cusp of true heat exhaustion. My life feels quite cushy in comparison when I sit down for a couple hours in AC with an iced coffee to command my computer to do work for me. And I dare say, compared to actual computer programmers, I don’t have that much code to write in the first place. What code I do write I find enjoyable, and the process helps me clarify my own thinking, much the way writing prose does. Getting a chair with optimal lumbar support is more relevant to my workflow than bothering with Claude Code.
Ultimately, the pro- and anti-LLM camps exaggerate the difference between quantitative social science with and without agentic tools. Various degrees of automation existed before. Code, like I already said, is just a set of instructions for routines that your computer will take care of for you. If you know what you’re doing, you can write very little code to make your computer automatically accomplish a great deal.
So who cares if you incorporate agentic LLMs into your academic workflow, or ask Claude to check your paper for grammatical errors or passive voice, or have an agent grab some data and write some analysis code? Who cares if you don’t adopt agentic LLMs, stick to the same old spell checker resources, or click a drop-down menu to download a dataset? While I admit to having some skepticism about agentic LLMs, and have been reluctant to adopt them into my own work, I use the online chat-based Claude on a regular basis for limited tasks, and I find it helpful. It takes care of boring grunt work for me. I write my own code, but I hate annotating it, so have Claude do it for me. I don’t always have the energy or time to proofread after a couple of hours of focused writing, so I let Claude read my work and come back with notes for things I need to fix. If someone uses it more intensively, that’s cool with me. If someone refuses to use these tools at all, I respect that, too. The final product is what matters most for professional success.
I’m with folks like Hollis Robbins that the drive toward sameness can be counterproductive. In her case, she means standardized teaching outcomes and course evaluations. I fear a drive toward stifling, sameness-enforcing standards for the adoption of LLMs in academia.
I wait, likely in vain, for a new equilibrium in my quantitative-social-science-corner of the academy where LLM users, abstainers, and those in the middle all can feel free to stop justifying their practices and no longer feel compelled to make others apologize for theirs. With these tools, researchers now have more choices. I like having choices. Don’t you?
Thanks for reading or listening! You can support Foreign Figures by liking, sharing, buying me a coffee, or becoming a free or paid subscriber.

