What belongs in a jury scoring sheet template — and what does not
A scoring sheet has one job: to turn several people's judgements into an order that can still be explained afterwards. Anything that does not serve that job makes the sheet worse. Five components are enough.
- Four or five criteria, no more. Each one gets a one-sentence definition — not “innovation”, but “how clearly does this differ from what is standard practice in the sector?”
- A short scale whose levels are described in words, not just numbered.
- A weighting per criterion, fixed before entries open.
- One mandatory comment box of two or three sentences. It disciplines the scoring and supplies material for feedback to entrants.
- Two exits: “I have a conflict of interest” and “I am not qualified to judge this”. Without them you force numbers that help nobody.
What does not belong is a criterion called “overall impression”: it restates the sum of the other criteria and quietly doubles the weight of what already counts. Also delete anything that confuses the quality of the submission with the quality of what was submitted. Layout, video editing and word count are indicators of budget, not achievement. If presentation matters to you, make it a separate, lightly weighted criterion and say so in the call for entries.
Why a 1–5 scale works better than 1–10
A ten-point scale looks more precise. It is not.
There is no shared meaning. Ask ten jurors what separates a 7 from an 8 and you get ten answers. With five levels each one can be described in a sentence, and everyone reads the same thing.
Nobody uses the whole range. On a 1–10 sheet the scores cluster almost entirely between 6 and 9; the bottom half stays empty because few people award a 2 to an entry someone paid to submit. Ten levels quietly become four — unlabelled ones.
It manufactures false precision. A 0.2-point lead on a ten-point scale looks like a result but is noise. If you base a decision on the second decimal place, you will have to defend it there, at the latest to the runner-up.
Spend the precision you save on describing the levels.
| Level | Label | How you recognise it |
|---|---|---|
| 5 | outstanding | Sets a new benchmark; I would cite this in the sector as an example |
| 4 | strong | Clearly above standard practice, with a visible contribution; minor weaknesses remain |
| 3 | solid | Competently done, but interchangeable |
| 2 | weak | Visible gaps in execution or in the evidence provided |
| 1 | inadequate | Misses the criterion, or the evidence is absent altogether |
Two objections are fair. In very large categories five levels produce many ties; the answer is not a finer scale but a first round that cuts the field to a shortlist (see the jury process). And if your jury parks everything in the middle, use an even 1–4 scale, which forces a direction.
Weighting: few round numbers, published in advance
Use steps of ten, add up to 100, and give no criterion less than 10 per cent. A weight of 17.5 per cent claims a precision of judgement that does not exist. A criterion worth 5 per cent can be deleted: it changes no outcome, but at 100 entries it costs every juror 100 decisions.
An example: impact 40 per cent, degree of innovation 30, execution 20, transferability 10. An entry scored 4 / 5 / 3 / 3 gives (4 × 0.4) + (5 × 0.3) + (3 × 0.2) + (3 × 0.1) = 1.6 + 1.5 + 0.6 + 0.3 = 4.0 out of 5. You can work that out on the phone, which is the point.
Two rules go with it. The weights are fixed before entries open and published with the call: people who know what counts submit better work. And they are not touched afterwards: a weighting changed once the scores are in is not fine-tuning, it is a correction of the result. In German-speaking markets, where trade associations and chambers run many long-established awards, publishing the weighting is now standard practice: members ask, and are entitled to.
Three criteria sets to start from
Treat these as starting points, not standards. Rewrite them in the language of your sector: criteria that entrants cannot restate in their own words will be answered wrongly.
Innovation or product award
- Novelty against current market practice — 30 %
- Demonstrated impact, backed by figures — 30 %
- Quality of execution — 20 %
- Scalability beyond the first adopter — 10 %
- Contribution to sustainability or resource use — 10 %
Project or campaign award
- Results against the documented starting position — 40 %
- Idea and creative execution — 20 %
- Craft quality of the delivery — 20 %
- Ratio of resources deployed to outcome — 10 %
- Transferability to other organisations — 10 %
Individual or rising-talent award
- Personal achievement in the period assessed — 40 %
- Effect beyond the person's own organisation — 20 %
- Leadership or role-model behaviour — 20 %
- Third-party evidence such as references and sources — 20 %
What all three share
- Five criteria at most, weights in steps of ten
- At least one criterion asks for evidence, not intent
- No “overall impression”, no criterion for layout
- A 1–5 scale with every level described
- Weights in the call for entries, not only in the jury briefing
Strict and lenient jurors: the fair average
Two jurors can use the same sheet and still pay in different currencies. One averages 3.2 points across everything she scores, the other 4.3. While both see the same entries, the difference cancels out. The moment you split the jury — and you will, once conflicts of interest or more than about 30 entries appear — the allocation becomes destiny.
| Entry | Scored by | Raw score | That juror's average | Deviation |
|---|---|---|---|---|
| A | strict juror | 4.0 | 3.2 | +0.8 |
| B | lenient juror | 4.2 | 4.3 | −0.1 |
On raw scores B wins, 4.2 to 4.0. Against each juror's own yardstick, A sits far above what that juror normally awards, while B sits below. The order reverses — with the same sheet.
Four ways to fix it
- Everyone scores everything. The cleanest fix and the only one needing no arithmetic. It holds to roughly 30 entries per category; beyond that your jurors' time runs out.
- Balanced allocation. Every entry receives the same number of scores, spread so each juror sees a comparable cross-section. It does not remove the difference in severity, but distributes it evenly.
- Score normalisation. Subtract each juror's own average from their raw scores, scale the result by how widely that juror spreads their marks, then map it back onto the original scale. Every juror is moved to the same zero point before averaging.
- Rank-based aggregation. Each juror puts their entries in order and only the ranks count. Immune to differences in severity, but it discards how far first place stood ahead of second.
When normalisation hurts: it assumes every juror saw a comparable cross-section. Someone who scores only five entries and happens to draw three excellent ones has an artificially high average subtracted — and their best entries lose. Rule of thumb: below roughly eight scores per person, stay with balanced allocation. And the method is fixed in advance: normalisation switched on after a first look at the ranking is no longer a method.
Keep any public vote out of this average. It measures reach rather than quality and belongs in its own category or as a clearly labelled additional score (more on public voting).
From the sheet to the process
The finished sheet is half the distance. Then the logistics: who receives which entry, who has a conflict, who has not submitted yet, and what unsuccessful entrants hear back.
Honestly: with 30 entries and five jurors scoring everything in one meeting, you do not need software. A spreadsheet with locked formulas will do, and the discussion is worth more than any automation. The effort tips over at three points — allocation, multiple rounds and feedback.
That is where Laureo comes in. Categories, criteria and scales are yours to configure in the Starter package (€ 1,790 net per award season); multiple jury rounds and allocation with conflict-of-interest handling sit in Professional (€ 3,490). Score normalisation, anonymised feedback PDFs and staged publication from shortlist to winner are part of the platform. Which function sits in which package is listed on the features page.
Scale anchors, criteria sets and weightings here come from practical award administration and are a starting point; your sector may need different criteria. Laureo is a product of State of Innovation GmbH, Vienna.
Frequently asked questions
How many criteria should a jury scoring sheet have?
Four or five. Each criterion needs a one-sentence definition and a weight of at least 10 per cent. More criteria do not add accuracy; they add judging time per entry and blur the sheet, because criteria that sit close together start measuring the same thing twice.
Is a 1-5 or 1-10 scale better for judging awards?
A 1 to 5 scale, in almost every case. With five levels you can describe each level in a sentence, so every juror means the same thing. On ten-point scales jurors use only 6 to 9 in practice, producing leads of a tenth of a point that nobody can justify.
How do you weight award judging criteria?
In steps of ten, adding up to 100, with no criterion below 10 per cent — for example 40/30/20/10. Fix the weights before entries open, publish them with the call for entries, and leave them untouched once scoring starts. Every placement then stays recalculable.
What is score normalisation in award judging?
A calculation that makes strict and lenient jurors comparable: each juror's own average is subtracted from their raw scores, the result is scaled by how widely that juror spreads marks, then mapped back onto the original scale. Useful from about eight scores per juror, and only if announced in advance.
Related reading
Planning an award right now?
In 20 minutes we walk through your process: entry, jury round, invoice.
Book a demo