06 Sep My Adventure in AI-Assisted “Writing” and Detection
Like most academics these days, I am very concerned about students misusing large language models (LLMs) like ChatGPT and ClaudeAI when they write papers for me. I don’t teach very often in my current position, but I do supervise masters theses. And there have been a number over the past few years that I suspected were written, in whole or at least in part, by AI.
I think most scholars would agree that asking an LLM to write part of or all of an academic article qualifies as research misconduct. Short of that, though, opinions seem to differ. Here is a list — courtesy of ChatGPT, of course — of academic tasks an LLM can do:
- Research assistance: identifying relevant topics, search terms, sources, or lines of inquiry.
- Source comprehension: explaining difficult concepts or summarizing sources selected by the student or scholar.
- Literature synthesis: comparing sources, identifying debates, or organizing the existing scholarship.
- Question development: refining the research question or suggesting possible hypotheses.
- Argument development: proposing claims, counterarguments, examples, or responses to objections.
- Structural planning: designing an outline and deciding how the argument should unfold.
- Editorial assistance: commenting on a written draft and suggesting revisions.
- Rewriting: substantially revising prose, organization, or argument.
- Passage generation: drafting discrete sentences, paragraphs, or sections for inclusion in the paper.
- Paper generation: writing the paper as a whole, with the student or scholar providing only a prompt, sources, or general instructions.
I’m on the traditionalist side of things, likely because LLMs didn’t exist for the first two decades of my academic career, I have no problem with #1; I often ask ChatGPT to create a thorough annotated bibliography for a subject I want to write about. I’m also okay with #2 (though using an LLM to summarise sources is a recipe for not understanding them) and #3. But I would consider #4 to #9 research misconduct and #8, #9, and #10 potentially plagiarism. (Because the LLM is trained on the work of other scholars.)
The problem, of course, is that it is very difficult to determine when a student or scholar has used an LLM in a way that I at least would consider research misconduct or plagiarism. Most AI-detection systems significantly underestimate AI-generated content while being relatively good at avoiding false positives, with Pangram being the best of the lot.
Reading a few studies of AI-detection, including the one linked to above, made me wonder how easy it would be to use an LLM to “write” an academic article that would be deemed 100% human-written even by a sophisticated AI-detector like Pangram. So, having just finished a long report for the Danish MoD and needing to decompress a bit, I tried an experiment. I came up with an idea for a short article (6,000 words or so) on a particular international criminal law issue raised by autonomous cyber weapons and tried to figure out the easiest way to achieve Pangram’s coveted 100% human-written score without having to write even one sentence from scratch.
I began by putting together, based on my own knowledge and information I found on the internet, codex instructions for ChatGPT to encourage it to write like a human — this human, in particular. I then wrote a 150-word abstract for the article that identified the general ICL issue I wanted to address, presented my basic argument, and provided a general list of the specific topics and questions I thought the article should cover. I then used a very simple prompt to turn ChatGPT loose: “I would like to see how you would use the abstract to write an academic article aimed at a specialty ICL journal. It should be approximately 6,000 words long, rely on the kind and number of sources appropriate for an article of that length, and cite sources like an academic would.”
19 minutes later, ChatGPT gave me the completed 6,400 word, 52 footnote article. The article was structured well, particularly in terms of its sections, and it did a good job making my argument and considering counter-arguments that have been made in the literature. I also checked the footnotes to make sure the LLM had not invented any sources or mis-cited real ones; it had not. But the legal analysis was a bit all over the place and lacked attention to detail, even if ChatGPT hadn’t made any important doctrinal mistakes. All in all, it was a reasonable first draft. Had I received it in an LLB or LLM course on international criminal law, I probably would have given it a B.
After reading the article, I started “revising.” The first thing I did was ask ChatGPT to create a hypothetical scenario that illustrated the central legal issue in the article, incorporate it into the introduction, and refer back to it as necessary in later sections. It came up with a very good scenario and rewrote the introduction and a few sections to incorporate it in less than five minutes. Impressive stuff. I then read the article again carefully and presented ChatGPT with notes about what I thought it could have done better. Some notes were general; others were more specific. Overall, I made 22 comments on the article.
Addressing my comments took ChatGPT another 17 minutes. The second draft was much stronger than the first — the logic of each section made more sense, the legal analysis was cleaner and more precise, the prose flowed more naturally. ChatGPT had even reversed the order of two sections in the way I would have, even though I had not suggested it. I am pretty sure I would have given this version of the article an A- in one of my ICL courses.
The process then proceeded iteratively: I read each new draft, offering general comments on the article and specific comments on each section while being careful not to tell ChatGPT how to fix the problems; ChatGPT revised; I offered further comments; and so on. I did three more full reads, eventually deeming the fifth draft of the article the final version.
I have little doubt that, substantively, draft five was more than strong enough to submit to a good journal — and would easily receive an A in one of my courses. But here my experiment broke down: despite going to great lengths in my codex instructions to encourage ChatGPT to write like a human, the “best” score a large chunk of text (approximately 500 words) achieved on Pangram was 75% AI-written. Points for Pangram!
At that point, I modified my experiment. I took one section of the ChatGPT-written article, approximately 700 words, and edited the text myself, adjusting the writing style to sound more like me. I did not write any sentences from scratch, but I did considerably revise a number of sentences: replacing words, switching around clauses, turning paragraphs into lists (“There are three reasons why X. First, Y. Second, Z”), etc.
In any case, manually editing one section of the article, which took me as long as it had taken ChatGPT to create five full drafts, did the trick: Pangram judged the revised section 100% human-written. Which, really, it wasn’t.
I would never submit the final article for publication, because I don’t consider it “my” work. In particular, ChatGPT carried out #4, #5, #6, #8, and #9 in the list above, and I don’t think a scholar (or a student) can ethically outsource any of those activities to an LLM. Unfortunately, if my Twitter feed is any indication (I know, I know), many scholars who should know better seem to disagree. Indeed, I am quite sure some of them would happily submit an article produced in the same way to an academic journal — all the while believing the work was theirs and they were doing nothing wrong.
That, to me, is the real danger LLMs pose for academia. I am good at research in my areas of expertise, and I write very quickly once I have my material organised and outlined. Working a few hours each day, I think I could have written the same article from scratch — hypothesis to finished product — in approximately six weeks. Outsourcing the article to ChatGPT, providing feedback on drafts, and doing just enough text editing for the final version to fool Pangram took me approximately eight hours. If scholars believe, or gradually come to believe, that submitting articles written as I “wrote” mine is a normal and acceptable academic practice, journals will be even more overwhelmed with submissions than they already are. And, worse, the very idea of scholarship will break down. (Think here of the University of Chicago business-school professor who has somehow “written” more than 250 articles this year — and it’s barely September.)
In the end, I think my experiment was — unfortunately — a success. That’s scary. And I have little doubt that the situation will become scarier still as LLMs become better at writing like a human. What the academy will do then, I have no idea.
NOTE: Kirsten Fisher has expressed a concern on Twitter that people will view this post as a how-to guide, not as a cautionary tale. I thought about that before I sat down to write. In the end, I decided that students and scholars who are okay with letting LLMs write for them won’t need my help to do it. I also think, given the blasé attitude toward LLM-produced scholarship I’ve seen among way too many scholars, that people need to understand that anyone can misuse an LLM, not just the tech-savvy younger generation. But feel free to weigh in on that issue here or on Twitter.

Sorry, the comment form is closed at this time.