Can AI Help Us Evaluate How We Write with AI: a proposal for a provisional manifesto for AI assisted writing
This article began with a conversation about AI and academic writing on LinkedIn. Dr. Susanne Friese who described asking the AI they regularly use to evaluate the way they worked together. I thought this was an excellent idea, partly because it turned AI-assisted writing back upon itself. Rather than asking whether a particular output was good, the question became: what kind of writing practice is developing through these repeated interactions, and what might that practice be doing to the writer?
I am in an unusually useful position from which to conduct such an experiment. I have been working with generative AI on academic writing for some time, across several articles at different stages of development and revision. Within ChatGPT, I also maintain an Academic Writing project containing drafts, conversations, instructions, source materials, editorial decisions and records of previous work.
Through repeated feeding (of my pre-AI writing), evaluation and correction, I have also developed what I call an academic house style. This is not simply a list of preferred spellings or grammatical conventions. It includes principles about argument, evidence, voice, theoretical framing, citation, structure and the preservation of enquiry in progress. It records recurring problems in AI-generated academic prose and instructs the system to avoid them. The house style is continually revised as new problems become visible.
There is then a substantial corpus through which the AI-assisted writing relationship could be examined. There is also a conceptual framework with which to examine it as I have recently completed work on an article on AI and academic writing. In the article I distinguish between inscription and orchestration. Inscription refers to the sentence-by-sentence activity through which writers have traditionally composed, revised and discovered what they think. Orchestration describes the increasingly complex work of coordinating prompts, outputs, sources, revisions, verification and integration.
The article argues that the central question is not simply whether AI has produced the words. It is where evaluative judgement occurs. If AI increasingly performs the work of noticing, diagnosing, comparing, selecting and revising, writers may retain control of the final text while offloading some of the activities through which their own judgement develops.
I therefore asked the AI to evaluate our working relationship through two connected questions.
First, to what extent had I preserved or offloaded evaluative judgement?
Second, had our elaborate house style resisted the homogenising tendencies of generative AI, or had it merely produced a more specialised and consistent form of homogenisation?
I explicitly asked for a critical evaluation rather than reassurance.
The result was not entirely comfortable. The AI concluded that I remained firmly in control of the research questions, conceptual direction, interpretation and final claims. I frequently rejected its diagnoses, challenged its formulations and identified occasions when fluent prose had distorted what I meant.
At the same time, I had offloaded a substantial amount of lower-level evaluative activity. I regularly asked the AI to identify structural problems, diagnose repetition, evaluate whether a section worked, compare alternative formulations, check whether reviewer concerns had been answered and decide whether an article appeared complete.
These are easily described as editorial tasks. But they also involve noticing, comparing, interpreting and judging. When AI becomes the default first reader, it does not merely offer solutions. It helps determine which problems become visible and which alternatives enter consideration.
The evaluation of homogenisation was equally mixed. The house style had clearly reduced generic AI prose. It had preserved important characteristics of my academic writing and repeatedly prevented unnecessary rewriting, artificial certainty and ceremonial uses of theory.
Yet it had not eliminated homogenisation. AI still exerted pressure towards orderly argument, balanced contrasts, smooth transitions, conceptual symmetry and an even academic register. More troublingly, there is a sense that the house style itself risked becoming a template through which very different research projects acquired the same rhetorical character.
What follows emerged from several further diagnostic and revision passes. I regard it as the beginning of a manifesto rather than a definitive protocol. It identifies principles for AI-assisted academic writing that seeks to preserve evaluative judgement and resist homogenisation.
A Manifesto for Judgement-Preserving AI-Assisted Academic Writing
AI-assisted academic writing is legitimate, but it is not intellectually neutral. Generative AI can contribute to language, structure, diagnosis, conceptual organisation and argumentative development. The boundary between assistance and collaboration is therefore neither clear nor stable.
The central question is not whether AI has participated in the writing. It is whether the academic writer continues to notice, compare, interpret, judge and justify.
1. Treat writing as an evaluative practice
Academic writing is not simply the production of fluent text. It involves recognising problems, testing claims, comparing alternatives, interpreting evidence and deciding what the work can legitimately say. These activities are part of academic judgement. They should not be routinely delegated simply because an AI system can perform plausible versions of them.
2. Make a human diagnosis first
Before asking AI to evaluate a draft, the writer should first identify what they believe is working, what remains uncertain and what may need revision. This diagnosis can be brief, but it should precede the AI’s assessment.
AI can then interrogate that diagnosis by identifying overlooked problems, questioning assumptions and proposing alternative interpretations. The writer must decide which elements of the resulting diagnosis are persuasive.
The sequence should be:
human diagnosis → AI critique → human judgement → revision
3. Diagnose before rewriting
AI should explain what it thinks is wrong before producing replacement prose. The writer can then evaluate the diagnosis independently of the fluency of the proposed solution.
A polished revision can make a questionable diagnosis appear correct. Separating diagnosis from rewriting prevents fluency from concealing the assumptions behind a change.
4. Do not confuse selection with full evaluative control
Choosing between AI-generated alternatives exercises judgement, but the AI has already shaped the available choices. Writers should sometimes formulate their own alternative, request genuinely different approaches, reject the options offered or ask what possibilities the AI’s framing has excluded.
The writer’s role must extend beyond selecting the most attractive item from a machine-generated menu.
5. Do not mistake fluency for quality
Generative AI is particularly good at producing orderly, grammatically controlled and conventionally academic prose. None of these qualities guarantees that a passage is accurate, original or intellectually worthwhile.
Fluent prose may conceal unsupported claims, unresolved contradictions, repetitive reasoning, imprecise concepts, weak connections between evidence and argument, or a degree of certainty the evidence does not support.
AI-generated prose should be evaluated for what it means and what it does, not simply for how professionally it sounds.
6. Preserve intellectual friction
Before using AI to revise a passage, the writer should identify any tensions, uncertainties, contradictions or unresolved questions that remain important to the enquiry. These should be treated as questions requiring attention rather than defects to be automatically removed.
When AI proposes a smoother or more coherent account, the writer should ask:
· What difficulty has this revision resolved?
· Was that difficulty genuinely resolved or merely written out of the text?
· Have competing ideas been made artificially compatible?
· Has uncertainty been replaced with confidence the evidence does not support?
· Does the revised account make the research process appear more orderly than it was?
Where a tension cannot yet be resolved, the writer should name and preserve it. AI can help clarify the disagreement, articulate competing interpretations or explore what would follow from different positions. The writer must decide whether synthesis is justified.
Preserving intellectual friction does not mean retaining confusion or resisting clarity. It means ensuring that revision clarifies a difficulty rather than making it disappear.
7. Develop a style without turning it into a template
Writers can develop explicit principles for working with AI: preferred terminology, prohibited clichés, characteristic rhythms, acceptable levels of formality and rules about preserving original wording. Such guidance can help resist generic AI prose.
These principles should not become a rigid template imposed on every project. Different subjects, methods and forms of enquiry may require different voices. The purpose of an academic writing practice is to remain responsive to the work, not to make every text consistently smooth.
8. Separate conceptual revision from stylistic editing
Conceptual and stylistic revision should normally occur as distinct operations.
During a conceptual pass, ask whether the argument is sound, whether the distinctions are meaningful, whether the evidence supports the claims and whether complexity has been simplified merely to produce coherence.
During a later stylistic pass, ask whether the meaning is clear, whether the prose is unnecessarily repetitive and whether terminology and tone are consistent.
AI should be instructed during a stylistic pass not to introduce new claims, concepts, evidence or argumentative relationships. Any substantive change should be returned to the writer for separate evaluation.
9. Read the sources
Academic sources must not be encountered only through AI-generated summaries, paraphrases or explanations. Writers must engage directly with the scholarship on which their arguments depend.
AI can assist with locating, organising, comparing and questioning sources, but it cannot substitute for reading them. Quotations must be checked, page references verified, contextual qualifications understood and AI interpretations tested against the original material.
This principle is non-negotiable because evaluative judgement depends upon having something more than the AI’s representation of a source against which to judge its claims.
10. Read the final text line by line
An AI audit is not a substitute for an authorial reading of the complete article. Before submission, the writer should read every sentence in sequence and ask:
· Do I understand and endorse this claim?
· Is this how I would express the idea?
· Does the sentence follow from what precedes it?
· Is the evidence sufficient?
· Has anything been made neater, stronger or more certain than it should be?
· Am I prepared to take responsibility for it?
Completion must remain an authorial decision.
11. Preserve a proportionate evidence trail
A complete transcript of every human–AI exchange may be impractical and may reveal little about where judgement occurred. Retaining only AI outputs creates an equally misleading record because it omits the questions, objections and rejected proposals that shaped them.
Writers should preserve a proportionate record of significant interventions. This might include version histories, major prompts, structural proposals and brief notes explaining why important recommendations were accepted, modified or rejected.
The purpose is not comprehensive surveillance. It is to keep consequential acts of judgement visible.
12. Describe AI involvement honestly
AI use should not be described as “language refinement” when the system has also contributed to diagnosis, structure, conceptual organisation or argumentative formulation.
An honest declaration should identify the kinds of contribution made by AI, explain how its outputs were evaluated and state clearly that the human author remains responsible for the sources, claims, interpretations and final text.
A Provisional Manifesto
This remains a provisional manifesto because stating these principles is considerably easier than practising them.
The first difficulty is self-discipline. “Read the sources” sounds obvious, but reading is slow and friction-heavy. It requires time and concentration. AI-generated summaries offer an immediate sense of comprehension, particularly when a writer is working quickly or across a large body of literature. The temptation is not necessarily to fabricate knowledge. It is to accept a plausible representation of a source as sufficient.
Yet direct engagement with the literature is where the writer acquires the basis for judging what the AI says about it. Without that independent encounter, verification can easily become superficial: checking that a source exists rather than assessing whether it has been understood.
The second unresolved problem is evidencing the process.
My own method has been to preserve versions of drafts, diagnostic reports, structural proposals and AI-generated revisions in separate documents. What I now realise is that this captures only one side of the collaboration. It records what the AI proposed but not necessarily the moments in which I questioned its diagnosis, rejected its language, reformulated the problem or insisted that a productive tension should remain unresolved.
Those interventions are precisely where evaluative judgement becomes visible.
Preserving every exchange is unlikely to be practical. Long transcripts can also obscure rather than illuminate the consequential decisions. What is needed is a lighter and more agile method: perhaps an automatically generated judgement record that identifies substantive turning points, records the original proposal and captures the writer’s reason for accepting, modifying or rejecting it.
I do not yet have a satisfactory method for doing this. It is one of the questions that this manifesto opens rather than answers.
What Would Honest Disclosure Look Like?
The evaluation also exposed a problem with conventional declarations of AI use. Many statements confine AI involvement to proofreading, language refinement or improvements in readability. Such descriptions may be technically convenient, but they are inaccurate when AI has also diagnosed weaknesses, proposed structures, generated alternatives or helped formulate conceptual distinctions.
The following declaration is my attempt to describe the practice more honestly, as it is, it is clearly too long and too wordy to be useful, but it serves as a starting point.
Declaration of generative AI use
Generative AI was used throughout the development and revision of this article as part of an iterative AI-assisted writing process. Its contributions included diagnostic review of drafts, structural exploration, comparison of alternative formulations, development and revision of passages, compression, identification of repetition, consistency checking and assistance with reference presentation. Its role was not confined to proofreading or language correction: at several points, the system proposed conceptual distinctions, argumentative sequences and possible ways of articulating the article’s claims.
The author originated the research problem and remained responsible for the article’s conceptual direction, interpretation of the literature, evaluation of evidence and final argument. AI-generated diagnoses and formulations were interrogated through repeated questioning, comparison, rejection and revision rather than accepted as authoritative. The developing manuscript was reviewed against a defined body of literature, and source-dependent claims, quotations and citations were subject to authorial verification. The final text was edited and approved by the author, who takes responsibility for its accuracy, integrity and conclusions.
This process is described as AI-assisted writing while acknowledging that the system’s contribution sometimes extended into activities conventionally understood as conceptual and compositional. The term “assisted” therefore identifies the author’s continuing responsibility for the work; it does not imply that AI participation was limited to mechanical or purely linguistic tasks.
Keeping Judgement Visible
AI-assisted writing does not require us to preserve a fiction of solitary, unassisted authorship. Nor should we pretend that generative AI is merely an advanced spelling and grammar checker when it has materially participated in the development of an argument.
The more important task is to keep judgement visible.
Who identified the problem? Who determined which evidence mattered? Who recognised that a fluent revision had removed an important tension? Who considered the alternatives that the AI did not offer? Who read the underlying sources? Who decided that the final text was accurate, defensible and complete?
“Human in the loop” is not, by itself, an adequate answer. A human can remain in the loop while exercising very little meaningful judgement. What matters is what the human continues to do there.
This manifesto is an attempt to specify those activities. It is also an invitation to test, challenge and extend them. If AI-assisted academic writing is becoming a durable scholarly practice, we need to move beyond arguments about whether it should occur and begin developing more demanding accounts of how it can be conducted without allowing convenience, fluency and optimisation to displace the formation and exercise of academic judgement.


Leave a Reply