AI & Technology

The Jagged Frontier: What the Evidence Says About AI Grant Writing

AI makes individual proposals better and every proposal more alike. Both findings are measured, and only one of them gets worse as adoption rises.

10 min readGuide: Evidence reviewUpdated August 2026

Eight hundred consultants walked into a field experiment. Some were given GPT-4, some were given GPT-4 plus training, and some were given nothing. Then they were handed two tasks that looked about equally hard.

On the first task, the AI groups outperformed the control group by 38% and 42.5%. On the second, they underperformed it — by 13 points and 24 points respectively. The group that had received training did worst of all.

That result is the most useful thing anyone has published about AI grant writing, and almost nobody in research administration talks about it. Not because the upside is fake — it is well measured and large — but because the shape of the downside is genuinely counterintuitive, and grant writing happens to be unusually exposed to it.

AI grant writing has a measured upside and a measured cliff

The study is Dell'Acqua and colleagues' Navigating the Jagged Technological Frontier, run with Boston Consulting Group across 758 knowledge workers and since published in Organization Science.

Inside the frontier — a creative product-ideation task that current models handle well — the gains were substantial. GPT-only participants improved 38% over the control group; those given GPT plus training improved 42.5%. The distribution matters as much as the average: less-skilled participants gained about 43%, while top performers gained about 17%. The tool compressed the gap between strong and weak performers.

Outside the frontier — an analytical task built with a subtle trap in the data — the same tool made people worse than having no tool at all. Not neutral. Worse. And the trained group's larger decline points at the mechanism the authors identified: participants tended to switch off their own judgment and follow what the model recommended.

The word doing the work here is jagged. The boundary between what AI does well and badly is not a smooth line running from easy to hard. Two tasks of apparently similar difficulty sat on opposite sides of it. You cannot locate the edge by asking how difficult something feels.

Why the cliff matters more than the gain for proposals

Grant writing is dense with tasks that sit on the wrong side of that boundary while looking like they sit on the right side.

Consider the asymmetry in how errors surface. A fabricated citation is embarrassing but cheap: it fails a thirty-second check, and the field has largely learned to look for it. That is a visible failure, and visible failures get caught.

Now consider a model confidently telling you that your preliminary data is sufficient to support an aim, or that a particular panel will find your approach novel, or that a work package is feasible in eighteen months. Each of those is a judgment call dressed as an answer. None of them fails a check, because there is no check. The error surfaces in a reviewer's comments a year later, or it never surfaces at all — it just quietly costs you the grant, and you attribute the loss to luck.

That is what makes proposal writing an uncomfortable fit for confident automation. The tasks where being wrong is invisible are exactly the tasks where "switching off your own judgment" carries no immediate penalty and no feedback signal. The red flags that reviewers have learned to spot in AI-assisted applications are the visible half of the problem. The invisible half is larger.

What "individually better, collectively worse" means when you are being ranked

The second study is the one that should change how you think about competitive advantage.

Doshi and Hauser, publishing in Science Advances, gave some writers story ideas generated by a large language model and left others to work unaided. The assisted stories were rated more creative, better written and more enjoyable — and, as in the BCG experiment, the benefit was largest for the less creative writers.

Then they measured the stories against each other. The AI-assisted stories were more similar to one another than the unaided ones were. The authors describe the result as a social dilemma: each writer is individually better off, while the collective pool of novel content narrows.

Transpose that to grant review and the implication is sharp. Grant panels do not score proposals against an absolute standard. They rank them against the other proposals in the pile. In a ranking system, a quality improvement that everyone receives is not an advantage — it is the new baseline. What survives as a genuine differentiator is whatever is not converging.

So the two findings interact in an awkward way. The tool reliably improves your individual output, and it reliably makes your output more like everyone else's, in a competition that is decided by distinctiveness. Both things are true simultaneously, and the second one gets worse as adoption rises rather than better.

This is not a speculative worry. It is the measured version of a warning the field has been issuing on instinct for two years — that heavy reliance produces more homogeneous applications — and it now has a study behind it rather than a hunch.

Who is actually adopting AI, and it is not who you would guess

The prevailing story about AI in academia is that well-resourced labs will pull further ahead. On at least one axis, the data points the other way.

Liu and colleagues analysed more than two million biomedical papers in PubMed Central from 2021 to 2024, estimating AI-assisted writing across the corpus. Their findings are specific: adoption grew roughly 400% in non-English-speaking countries against 183% in English-speaking ones, and it was highest among less established scientists — those with fewer publications, lower citation counts, and positions at lower-ranked institutions.

The associated effect was a modest productivity increase and a measurable narrowing of the publication gap between researchers in English-speaking and non-English-speaking countries.

Anyone who has watched a strong project lose to a fluently written weaker one will recognise what is being dismantled here. The penalty for writing science in a second language is real, documented, and has nothing to do with the quality of the underlying work. A tool that reduces it is doing something straightforwardly good, and the people reaching for it first are the people it helps most.

Hold that beside the homogenisation finding without resolving the tension too quickly. The same technology is lowering an unfair barrier and narrowing the diversity of what gets submitted. Both effects are real. A story that only tells you one of them is selling something.

Mapping your own frontier for AI grant writing

The practical question is not whether to use these tools. It is where the edge runs in your own workflow. The test that holds up: can you detect the error yourself, cheaply, without outside help? If yes, you are inside the frontier. If no, you are outside it, and the model's confidence is worth nothing.

Inside the frontier Outside the frontier
Sweeping literature and prior art for things you missed Judging whether your idea is actually novel
Checking a draft against eligibility and formatting rules in the call text Deciding what this specific panel will reward
Tightening prose whose logic you have already worked out Assessing whether your preliminary data carries an aim
Translating a specialist section for a general panel Feasibility and timeline judgments
Generating alternative structures for a section you are stuck on Anything depending on unpublished context — lab politics, a programme officer's steer, why the last round failed
First-pass arithmetic and consistency checks on a budget Deciding what to cut when the budget does not fit

The pattern in the left column is that verification is cheap and failure is visible. The pattern in the right column is that the output looks equally polished whether it is right or wrong.

One consequence worth naming: the left column is mostly the work that used to consume your evenings, and the right column is mostly the work that determines whether you get funded. That is a good trade, but only if you actually reallocate the recovered time to the right column instead of submitting more applications. The discipline of building a reusable context layer helps here, precisely because it front-loads the judgment calls into something you have already reasoned about once.

The move that follows from the evidence: critic, not author

If the danger sits in tasks where failure is invisible, then the fix is not better prompting. It is changing the direction the work flows.

Ask a model to generate your novelty statement and you are outside the frontier: you have no independent way to check whether what comes back is genuinely distinctive or a confident average of everything similar in its training data. Ask the same model to attack your novelty statement — list the three nearest existing approaches, name the reviewer objection you have not answered, argue that this is incremental — and you have moved into the visible-failure regime. You can evaluate every objection it raises, because you are the one who knows the field. Bad objections are obvious. Good ones are gifts.

The asymmetry is worth stating plainly. Generation asks the model for something you cannot verify. Critique asks it for something you can. Same tool, same session, opposite risk profile.

In practice this means using it hardest on the sections you are most confident about, which feels backwards. Confidence is where blind spots live. A specific aims page you have rewritten nine times is a page you can no longer read as a stranger, and an adversarial reader that never gets bored is genuinely useful there — far more useful than it is at drafting the page in the first place.

A worked example, since the distinction is easy to nod along to and hard to apply. Take a feasibility argument for a two-year work package:

  • Outside the frontier: "Is this timeline realistic?" You cannot check the answer. The model has no knowledge of your equipment queue, your technician's notice period, or how long the ethics approval took last time.
  • Inside the frontier: "Here is my timeline and the four assumptions it rests on. Which assumption, if wrong, breaks the schedule worst?" Now it is doing dependency analysis on information you supplied, and you can audit every step of the reasoning.

The second prompt is more work to write. That is the point — the work of specifying your own assumptions is the thinking, and it is not delegable. What you get back is a stress test rather than an answer.

This also explains why the "AI wrote my proposal" failure mode is not primarily an integrity problem. It is a feedback problem. An author who never receives a real objection never discovers the weakness, and a model asked to generate rather than critique is optimised to agree with the premise of the question.

What this changes

Three things follow from the evidence, and none of them is "use AI more" or "use AI less."

First, treat the model's confidence as uninformative. It is uncorrelated with whether you are inside or outside the frontier, which is exactly why the trained group in the BCG study fell furthest. Training raised trust without moving the boundary.

Second, protect the parts of a proposal that make it identifiably yours — the framing of the problem, the reason this team and not another, the specific intellectual bet. Those are the sections where convergence costs you most, and they are also the sections most tempting to hand over because they are the hardest to write.

Third, spend the time you recover on judgment rather than volume. If a tool saves you thirty hours and you spend those hours on two more applications, you have converted a quality gain into a quantity gain in a system that ranks on quality. Support staff and institutional workflows tend to make this mistake structurally, because throughput is easier to measure than distinctiveness.

The honest summary of the evidence on AI grant writing is that it works, that it works best for the people who have been least well served by the status quo, and that its aggregate effect is to make fluency abundant and distinctiveness scarce. Fluency used to be a differentiator. It is becoming table stakes. What you have that the model does not — the unpublished context, the judgment about what matters, the reason this problem is worth six years of your life — is now the whole of your advantage, and it is the one thing no tool will supply.

EG

Founder & CEO, Proposia.ai

PhD researcher and Associate Professor in Computer Science, working at the intersection of algorithm design, applied mathematics, and machine learning. With Proposia.ai, I aim to transform research ideas into scalable AI solutions that support innovation and discovery.