Not Only Luck.
An essay by John Melonakos

Using the Full Specturm for Ratings or Grades

Over the past several years, I participated in the review of NIH SBIR grants. The scale for grading the grants is from 1 (best) to 9 (worst). Traditionally, reviewers would default to a 3, with many getting a 2 or a 4, and a few getting a 1. Few would receive above a 4.

All the scores were bunched up into the 1-4 squished portion of the grading scale, leaving 5-9 unoccupied. The resolution of differentiation among the grants was corresponding diminished. As a reviewer, you felt you had to conform to this system bias in order for your scores to have impact.

Grades in high school are similarly squished with 75-100 being where all the action occurs and 0-75 remaining unoccupied territory. Grade inflation in educational institutions makes this even worse.

This week I have been reviewing another batch of grants. This time all the reviewers were instructed to use more of the scale. An average grant should get a 5. A poorer than average grant should get above a 5. A better than average grant should get below a 5. It’s a recalibration of the scale in an effort to provide more meaningful resolution of the grant spread and a more accurate correlation between the meaning of the numbers on the scale and the scores received by the proposals.

In a few of my electrical engineering courses at BYU, professors would treat the scale from 0 to 100 the same way, the median grade should be close to a 50. The professors would target a level of difficulty that would yield as close to a uniform distribution outcome as possible.

The uniform distribution is the best target distribution if you want to most effectively use the resolution available in any rating system or scale.

Rating scales for customer or employee satisfaction, surveys to clients, or other metric gathering should be designed with effective use of the full spectrum in mind. It does no good to ask a scale-based question wherein everyone picks the “10” and no valuable information is actually generated.

Using the full scale can be a painful process, but next time you design a rating system, give some thought to targeting a uniform distribution!

What are your thoughts on rating scales?

 

Related articles