My Adventure in AI-Assisted “Writing” and Detection

My Adventure in AI-Assisted “Writing” and Detection

Like most academics these days, I am very concerned about students misusing large language models (LLMs) like ChatGPT and ClaudeAI when they write papers for me. I don’t teach very often in my current position, but I do supervise masters theses. And there have been a number over the past few years that I suspected were written, in whole or at least in part, by AI.

I think most scholars would agree that asking an LLM to write part of or all of an academic article qualifies as research misconduct. Short of that, though, opinions seem to differ. Here is a list — courtesy of ChatGPT, of course — of academic tasks an LLM can do:

  1. Research assistance: identifying relevant topics, search terms, sources, or lines of inquiry.
  2. Source comprehension: explaining difficult concepts or summarizing sources selected by the student or scholar.
  3. Literature synthesis: comparing sources, identifying debates, or organizing the existing scholarship.
  4. Question development: refining the research question or suggesting possible hypotheses.
  5. Argument development: proposing claims, counterarguments, examples, or responses to objections.
  6. Structural planning: designing an outline and deciding how the argument should unfold.
  7. Editorial assistance: commenting on a written draft and suggesting revisions.
  8. Rewriting: substantially revising prose, organization, or argument.
  9. Passage generation: drafting discrete sentences, paragraphs, or sections for inclusion in the paper.
  10. Paper generation: writing the paper as a whole, with the student or scholar providing only a prompt, sources, or general instructions.

I’m on the traditionalist side of things, likely because LLMs didn’t exist for the first two decades of my academic career, I have no problem with #1; I regularly ask ChatGPT to create thorough annotated bibliographies for subjects I want to write about. I’m also okay with #2 (though using an LLM to summarise sources is a recipe for not understanding them) and #3. But I would consider #4 to #9 research misconduct and #8 and #9 (and obviously #10) potentially plagiarism. (Because the LLM is trained on other scholars’ writing.)

The problem, of course, is that it is very difficult to determine when a student or scholar has used an LLM in a way that I at least would consider research misconduct or plagiarism. Most AI-detection systems significantly underestimate AI-generated content while being relatively good at avoiding false positives, with Pangram being the best of the lot.

Reading a few studies of AI-detection, including the one linked to above, made me wonder how easy it would be to use an LLM to “write” an academic article that would be deemed 100% human-written even by a sophisticated AI-detector like Pangram. So, having just finished a long report for the Danish MoD and needing to decompress a bit, I tried an experiment. I came up with an idea for a short article (6,000 words or so) on a particular international criminal law issue raised by autonomous cyber weapons and tried to figure out the easiest way to achieve Pangram’s coveted 100% human-written score without having to do any actual writing myself.

I began by putting together, based on my own knowledge and information I found on the internet, codex instructions for ChatGPT 5.6 Sol in the Extra High setting. Here they are:

I am a professor of international law and security. My primary research areas include the jus ad bellum and jus in bello (especially their history and development), new weapons technology, international criminal law, and ecocide. When you do academic research for me, it is important to rely only on high-quality academic and popular sources that I can check. I prefer peer-reviewed articles, monographs from leading academic presses (like OUP and Belknap) and scholarly trade presses (like Verso), and primary legal documents, although articles in quality non-peer-reviewed journals and magazines aimed at an educated lay audience are also fine. For historical research, archival and contemporaneous sources are especially valuable. You should always provide a range of sources that include different viewpoints on what I research, and you should try to avoid being too US or even European-centric.

In terms of writing style, try to use prose that mirror my academic voice, which you can identify by reading some of my published work. I like to use lists when making arguments (first, second, third…) and to break sections into subsections. No antithesis. No corrective negation. Minimal paragraph pinning. Minimal parataxis. No summary beats. No rhetorical crutches. No negative parallelisms. No negative anaphoras. No contrasting pairs. No rule of three. Minimise em dashes even though I love them. No throat-clearing openers. Minimal landing sentences. No setup/payoff constructions. No parallel sentence structures within a paragraph. Vary sentence length unpredictably. No stacked noun phrases. Minimal filler intensifiers (genuinely, really, truly, actually). No corporate-speak verbs (leverage, underscore, reflect). No nominalization. No hedging qualifiers. No performed enthusiasm.

After inputting the codex instructions, I wrote a 150-word abstract for the article that identified the general ICL issue I wanted to address, presented my basic argument, and provided a general list of the specific topics and questions I thought the article should cover. I then used a very simple prompt to turn ChatGPT loose: “I would like to see how you would use the abstract to write an academic article aimed at a specialty ICL journal. It should be approximately 6,000 words long, rely on the kind and number of sources appropriate for an article of that length, and cite sources like an academic would.”

19 minutes later, ChatGPT gave me the completed 6,400 word, 52 footnote article. The article was structured well, particularly in terms of its sections, and it did a good job making my argument and considering counter-arguments that have been made in the literature. I also checked the footnotes to make sure the LLM had not invented any sources or mis-cited real ones; it had not. But the legal analysis was a bit all over the place and lacked attention to detail, even if ChatGPT hadn’t made any important doctrinal mistakes. All in all, it was a reasonable first draft. Had I received it in an LLB or LLM course on international criminal law, I would have probably given it a B.

After reading the article, I started “revising.” The first thing I did was ask ChatGPT to create a hypothetical scenario that illustrated the central legal issue in the article, incorporate it into the introduction, and refer back to it as necessary in later sections. It came up with a very good one and rewrote the introduction and a few sections to incorporate it in less than five minutes. Impressive stuff. I then read the article again carefully and presented ChatGPT with notes about what I thought it could have done better. Some notes were general, like “Tighten Section II” or “don’t include normative claims in the substantive sections. Collect them and present them in a logical way in the conclusion.” Others were more specific, such as “explain the difference between recklessness and dolus eventualis with greater precision and provide at least one example where the difference would matter” or “I think you should describe autonomy in more detail using Scholar X’s work instead of Scholar Y’s.” Overall, I made 22 comments on the article.

Addressing my comments took ChatGPT about 17 minutes. The second draft was much stronger than the first — the logic of each section made more sense, the legal analysis was cleaner and more precise, the prose flowed more naturally. ChatGPT had even reversed the order of two sections in a way that I would have, even though I had not suggested it. I am pretty sure I would have given this version of the article an A- had a student submitted it in one of my ICL courses.

The process then proceeded iteratively: I read each new draft, offering general comments on the article and specific comments on each section while being careful not to tell ChatGPT how to fix the problems; ChatGPT revised; I offered further comments; and so on. I did three more full reads, deeming the fifth draft of the article the final version.

I have little doubt that, in terms of its substance, draft five was more than strong enough to submit to a good journal — and would easily receive an A in one of my courses. But here my experiment broke down: despite going to great lengths in my codex instructions to encourage ChatGPT to write like a human, the “best” score a large chunk of text (approximately 500 words) achieved on Pangram was 75% AI-written. Points for Pangram!

At that point, I modified my experiment. I took one section of the ChatGPT-written article, approximately 700 words, and edited the text myself, adjusting the writing style to sound more like me. I did not write any sentences from scratch, but I did considerably revise a number of sentences: replacing words, switching around clauses, turning paragraphs into lists (“There are three reasons why X. First, Y. Second, Z”), etc.

In any case, the manual editing of one section of the article, which took me as long as it had taken ChatGPT to create five full drafts, did the trick: Pangram judged the revised section 100% human-written. Which, really, it wasn’t.

I would never submit the final article for publication, because I don’t consider it “my” work. In particular, ChatGPT engaged in #4, #5, #6, #8, and #9 in the list above — and arguably even #10 — and I don’t think a scholar (or a student) can ethically outsource any of those activities to an LLM. But if my Twitter feed is any indication (I know, I know…), many scholars who should know better seem to disagree. Indeed, I am quite sure some of them would happily submit an article produced in the same way to an academic journal — all the while believing that the work was theirs and they were doing nothing wrong.

That, to me, is the real danger LLMs pose for academia. I am good at research in my areas of expertise, and I write very quickly once I have my material organised and outlined. Working a few hours each day, I think I could have written the same article from scratch — hypothesis to finished product — in approximately six weeks. Outsourcing the article to ChatGPT, providing feedback on drafts, and doing just enough text editing for the final version to fool Pangram took me approximately eight hours. If scholars believe, or gradually come to believe, that submitting articles written as I “wrote” mine is a normal and acceptable academic practice, journals will be even more overwhelmed with submissions than they already are. And, worse, the very idea of scholarship will break down. (Think here of the University of Chicago business-school professor who has somehow “written” more than 250 articles this year — and it’s barely September.)

In the end, I think my experiment was — unfortunately — a success. That’s scary. And I have little doubt that the situation will become scarier still as LLMs become better at writing like a human. What the academy will do then, I have no idea.

NOTE: I would love to hear what others think. Please comment on this post here or on Twitter!

Print Friendly, PDF & Email
Topics
Artificial Intelligence, Featured, International Criminal Law, International Law, Legal education, Technology

Leave a Reply

Please Login to comment
avatar
  Subscribe  
Notify of