Psychology & Mindset

Grant Peer Review Reliability: Why Every Source Quotes a Different Number

Zero, 0.26, 0.5, 59% — all published, all defensible. The disagreement is in the units, not the data, and the distinction decides what you do after a rejection.

9 min readGuide: Evidence reviewUpdated August 2026

Go looking for how reliable grant peer review actually is and you will find, in reputable places, all of the following: zero. 0.15. 0.26. 0.34. 0.50. Fifty-nine percent.

That looks like a field unable to measure its own central process. It isn't. Every one of those figures is defensible, and they disagree because they are measuring different things — different units, different populations, different slices of the score range. The confusion is not in the data. It is in the summarising.

This matters beyond pedantry, because the number you believe determines what you do after a rejection. If grant peer review reliability is zero, resubmitting the same proposal is rational and polishing it is superstition. If it is 0.5, the reverse. Getting this right is worth more than another week of revision.

Grant peer review reliability: the numbers and what each one measured

Figure Study What was actually measured
ICC ≈ 0 Pier et al., PNAS 2018 Agreement between independent reviewers scoring the same R01 application
ICC = 0.259 Mutz, Bornmann & Daniel, PLOS ONE 2012 Reliability of a single reviewer's rating, Austrian Science Fund
ICC = 0.495 Same study, same data Reliability of the mean rating across 2.82 reviewers
ICC 0.183–0.319 Same study, by field Biosciences at the low end, humanities at the high end
59% of funded grants Graves, Barnett & Clarke, BMJ 2011 Share of 620 funded NHMRC grants that would sometimes miss out once score variability is modelled

Notice that rows two and three come from the same dataset and differ by a factor of nearly two. Nothing about reviewer behaviour changes between them. Only the unit of analysis does.

One reviewer or the whole panel? This resolves most of the confusion

The gap between 0.259 and 0.495 is not a discrepancy. It is arithmetic.

Mutz and colleagues analysed Austrian Science Fund proposals and reported both: a single reviewer's rating carries an ICC of about 0.26, while the average rating across the 2.82 reviewers a proposal typically received carries an ICC of about 0.50. Averaging independent noisy judgments cancels some of the noise. This is the oldest result in psychometrics, and it is the reason panels exist at all.

So two statements that sound contradictory are both true. Individual reviewers agree with each other poorly. The panel's aggregate judgment is substantially more reliable than any individual in it.

This is where the popular summary goes wrong. "Reviewers don't agree, so grant review is a coin flip" takes a single-reviewer statistic and applies it to a panel decision. The panel is the thing that decides, and it is measurably better than its parts. Not good — 0.5 is not good — but not a coin flip either.

The same logic runs the other way, and it is useful when you read a rejection. A criticism made by one reviewer is weak evidence about your proposal. The same criticism arriving independently from three reviewers is strong evidence, and it is strong because baseline agreement is so low. When agreement is rare, convergence is informative. That inversion is the single most practical thing to take from this literature.

Why Pier et al. found zero, and what to take from it

The study that produces the most alarming number deserves to be read carefully rather than quoted.

Pier and colleagues recruited 43 experienced oncology reviewers, constituted them into 16 simulated NIH study sections, and had them evaluate 25 real R01 applications. The reported intraclass correlations were 0 for overall ratings, 0 for strengths, and 0.017 for weaknesses. Reviewers scoring the same application resembled each other about as much as reviewers scoring different applications.

That is a genuinely important result and it should not be softened. It should also be read with its limits visible: a single field, a modest number of applications, simulated panels without real funding consequences, and — this one matters most — the applications used had already been funded.

That last detail is not a debunking. It is the mechanism, and it explains almost everything about why these numbers scatter.

Range restriction: why reliability looks worst exactly where you care most

Correlation depends on variance. Remove the variance and the correlation collapses, regardless of how good the raters are.

If you assemble a set of applications that all previously cleared the funding bar, you have deliberately discarded the weak ones. What remains is a narrow band of strong proposals plus the ordinary noise of human judgment. A near-zero correlation in that band is close to a statistical inevitability, not a finding about reviewer competence.

The practical translation is uncomfortable and, I think, correct. Peer review is probably decent at separating the strongest third from the weakest third, and close to useless at separating the proposal ranked eighteenth from the one ranked twenty-second. Both claims come from the same evidence base.

The problem is that funding decisions are made almost entirely in the second regime. Success rates in the teens mean the interesting action all happens inside the compressed band where the signal has already been spent. That is a design problem in how funding lines are drawn, not evidence that reviewers are bad at reading science.

It also puts the Graves study in context. Finding that 59% of 620 funded NHMRC grants would sometimes go unfunded under modelled score variability is a statement about proposals near a cutoff, which is where nearly all funded proposals sit when money is scarce. It is not a claim that 59% of funding decisions are arbitrary in any deeper sense.

The field effect nobody expects

One result from the Austrian data is worth sitting with because it runs against intuition: single-rater reliability was lowest in the biosciences (0.183) and highest in the humanities (0.319).

The fields with the most methodological consensus produced the least agreement between reviewers. That is the opposite of what most researchers would predict.

The likeliest explanation is range restriction again, operating at the level of the discipline. Where methods are standardised and training is uniform, a large share of submitted proposals are technically competent, so the quality distribution compresses and reviewers are left discriminating on taste. Where traditions are more varied, proposals differ more, and there is more real variance for reviewers to agree about.

I would treat that as a plausible reading rather than a settled finding — the study identifies the pattern, not the cause. But if it holds, it means "we have rigorous shared standards" and "our panels will disagree a lot" are compatible, even connected.

Some of it is fixable, which tells you what kind of problem it is

Here is the finding that changed how I read this whole literature.

Sattler, McKnight, Naney and Mathis randomly assigned 75 public health professors to watch an eleven-minute training video before reviewing, or not. The video did nothing clever — it explained what each value on the NIH rating scale means, why scores matter, and that the review criteria are worth reading carefully.

Inter-rater reliability went from an ICC of 0.61 without the video to 0.89 with it. Scoring accuracy went from 35% correct to 74%.

Two caveats before anyone gets excited. That baseline of 0.61 is far above what the field studies report, because this was a controlled task with explicit criteria and a handful of proposals rather than a real study section grinding through a pile at eleven at night. And 0.89 is not a number any funder should expect to reproduce at scale.

But the direction is the interesting part, and it reframes the whole question. If eleven minutes of instruction about how to use the scale nearly halves the disagreement, then a substantial share of what we have been calling "reviewers disagree about science" is actually "reviewers are using the same numbers to mean different things." That is a calibration failure, not an irreducible clash of expert judgment. Calibration is a solved problem in other fields.

It also hands applicants something concrete. If part of the variance comes from reviewers interpolating between vague criteria and your text, then anything that removes the need to interpolate is doing real work on your behalf. Mapping your sections explicitly onto the published review criteria — using the funder's own language for the funder's own scoring dimensions — is not cosmetic compliance. It reduces the number of judgment calls a tired reviewer has to invent, and every invented judgment call is a draw from a distribution you would rather not sample.

That is a rare thing in this literature: a lever the applicant actually controls, supported by a mechanism rather than a vibe.

What grant peer review reliability means for your next decision

Four things follow, and they are more specific than "don't take it personally."

A borderline rejection carries very little information about quality. In Horizon Europe's first two years, 71% of proposals scoring above the quality threshold went unfunded for budget reasons alone, with oversubscription running at roughly 4.7 times available budget. Above-threshold rejection is the default outcome of an oversubscribed programme, not a verdict. Two proposals scoring 14.4 and 14.6 out of 15 are not distinguishable by any measurement the panel possesses.

Resubmission is mechanically a fresh draw. A new round usually means a different set of reviewers, and given what the reliability data shows, that is a real effect rather than hopeful thinking. This is why resubmitting rather than repackaging is generally the higher-expected-value move after a near miss — the odds genuinely improve, and not only because the text got better.

Weight convergent criticism heavily and divergent criticism lightly. If three reviewers independently flag your feasibility section, that is signal. If one reviewer dislikes your framing and two do not mention it, you are looking at noise, and rewriting to satisfy the outlier can easily make the proposal worse for the next panel. Diagnosing which kind of rejection you received is a separate skill worth learning properly.

Do not run a post-mortem on a single data point. One rejection from one panel, at an ICC of 0.26 per reviewer, is thin evidence about anything. Patterns across three submissions are worth analysing; a single set of comments mostly tells you about the reviewers you drew. That is also why a structured post-mortem beats an emotional one — it forces you to separate the recurring from the incidental.

What this evidence does not license

Two readings are tempting and both are wrong.

The cynical reading — the system is random, so proposal quality is irrelevant — does not follow. Low reliability inside a restricted range is fully compatible with peer review doing real work across the full range. Nobody has shown that a weak proposal and a strong one have similar prospects. They demonstrably do not.

The naive reading — reviewers are experts, so scores reflect merit — does not survive contact with Pier et al. At the margin, which is where your application probably sits, the score is substantially a property of which reviewers you got.

The defensible position is narrower and more useful than either. Grant peer review reliability is adequate for coarse sorting and inadequate for fine ranking, and the funding line almost always falls in the fine-ranking region. Write well enough to clear the coarse sort decisively, because that part is real and you control it. Then treat everything that happens in the last few percentiles as a lottery you have bought a good ticket for — and buy more tickets, because that is the only lever the evidence actually supports.

EG

Founder & CEO, Proposia.ai

PhD researcher and Associate Professor in Computer Science, working at the intersection of algorithm design, applied mathematics, and machine learning. With Proposia.ai, I aim to transform research ideas into scalable AI solutions that support innovation and discovery.