Curve Exam Scores

Skip to main content
How do I...?
Categories
< All Topics
Print

How Do I Curve Exam Scores?

What is Curving?

Curving refers broadly to adjusting students’ raw assessment scores after an assessment has been administered. However, the term is used to describe several different practices, not all of which are equally appropriate.

Traditional curving adjusts grades based on how students perform relative to one another. Other approaches adjust scores to address a problem with an assessment, an unexpectedly difficult examination, or a mismatch between observed performance and the intended performance standard.

At NECO, we strive to employ criterion-referenced approaches: students are evaluated according to whether they have demonstrated established knowledge, skills, and competencies rather than according to how they perform relative to their classmates. For this reason, routine curving is generally discouraged.

When is Curving Appropriate?

There are circumstances in which adjusting assessment scores is appropriate and defensible. The important question is why the adjustment is being made.

A score adjustment is most defensible when there is evidence that the original scores do not accurately represent student achievement because of a problem with the assessment, its administration, or the calibration of its difficulty.

In these situations, the goal is not to create a desirable grade distribution. The goal is to produce scores that more accurately represent student performance relative to the intended standard.

Defensible Curving Methodologies

Remove or Rescore Flawed Items

This should generally be the first approach considered when unexpectedly poor performance can be traced to specific assessment items. An item may warrant removal or rescoring when there is evidence that it:

  • Contains no defensible correct answer
  • Contains more than one defensible correct answer
  • Is ambiguous or misleading
  • Includes a factual or technical error
  • Assesses content that was not appropriately represented in the course or learning objectives
  • Was affected by a display, technology, or administration problem

Depending on the problem, an item may be removed from scoring, multiple responses may be accepted, or credit may be awarded to affected students.

This is often preferable to adjusting the entire examination because the correction is targeted to the identified source of measurement error.

Apply a Uniform Score Adjustment

A fixed number of points or percentage points may be added to all students’ scores when there is a defensible reason to conclude that the examination as a whole was more difficult than intended.

For example, if review suggests that an examination was systematically more difficult than comparable examinations or exceeded the intended level of cognitive demand, a uniform adjustment may help recalibrate scores. Because every student receives the same adjustment, relative differences among students are preserved.

However, the size of the adjustment should be based on a reasoned evaluation of the assessment rather than on the number of points necessary to produce a desired pass rate or average.

Apply a Transformation to Account for Assessment Difficulty

Mathematical transformations, such as linear adjustments or other score conversions, can sometimes be used when assessment difficulty differs meaningfully from what was intended.

These approaches may be appropriate when there is sufficient assessment data to support the transformation and a clear rationale for how the transformed scores should be interpreted.

Because transformations can alter the relationship between raw performance and reported grades, they should be used cautiously and preferably with consultation from the TLC.

Curving Practices to Avoid

Some common approaches are difficult to justify within a criterion-referenced assessment system.

Grading on a Normal Distribution

Predetermining that a certain proportion of students will receive each grade makes performance explicitly dependent on the performance of other students. For example, assigning the top 10% of students an A and the bottom 10% a failing grade is a norm-referenced approach. It does not establish whether either group actually met the intended learning expectations.

This approach is generally inconsistent with criterion-referenced assessment.

Curving to Achieve a Desired Class Average

Adding enough points to produce a predetermined class mean, such as 80%, is also problematic when the target is based solely on what the grade distribution is expected to look like. A class average is a description of group performance. It is not, by itself, evidence that an assessment was too difficult or that scores are inaccurate.

Curving to Produce a Desired Pass Rate

Similarly, an unexpectedly high failure rate should trigger investigation, not an automatic curve. A high failure rate may reveal a problem with the examination, but it may also reflect genuine deficiencies in student learning. Adjusting scores simply until an acceptable percentage of students pass risks obscuring that distinction.

Using the Highest Student Score as 100%

Some approaches calculate an adjustment based on the highest-performing student’s score, such as adding enough points to make the highest score equal 100%. This assumes that at least one student should have achieved a perfect score. There is generally no criterion-referenced reason that this must be true.

The strongest student’s performance is not itself evidence of the appropriate standard for everyone else’s performance.

Before Curving an Examination

Before making a score adjustment, consider:

  • Was there a problem with specific items?

Review item statistics and examine poorly performing items for ambiguity, errors, miskeying, or other flaws.

  • Was the assessment appropriately aligned?

Confirm that the examination reflected the stated learning objectives, instructional emphasis, and expected cognitive level.

  • Was the examination unusually difficult?

Compare performance and assessment characteristics with previous examinations or other relevant evidence when available.

  • Does the performance pattern reflect an assessment problem or a learning problem?

Poor performance does not necessarily mean an examination was unfair or overly difficult.

  • What interpretation should the adjusted score have?

The adjustment should make scores more representative of achievement relative to the intended standard, not simply make the grade distribution more desirable.

  • Can the adjustment be explained and defended?

Faculty should be able to articulate why the adjustment was necessary, how the method was selected, and why it was applied consistently and fairly.

Need Help Evaluating an Exam?

The TLC can assist faculty with reviewing item and examination performance, identifying potentially problematic items, interpreting assessment data, and determining whether a score adjustment is appropriate.

Contact the TLC before applying a substantial or unusual score adjustment if you would like assistance determining the most defensible approach.

Was this article helpful?
0 out of 5 stars
5 Stars 0%
4 Stars 0%
3 Stars 0%
2 Stars 0%
1 Stars 0%
5
Please Share Your Feedback
How Can We Improve This Article?
Table of Contents