Education: Joining the Dots
The answer lives in this podcast

Answer extracted from the Education: Joining the Dots podcast — listen to the full episode below.

🎧 Listen to the episode on Listenly

Why internally designed tests fail to compare student performance fairly

Internally designed tests are inherently inconsistent in difficulty, making them poor tools for comparing performance across different students, time periods, schools, or subjects. Standardized tests solve this problem by converting raw scores onto a scaled distribution, enabling fair comparison against national benchmarks and meaningful tracking over time.

The core problem: no common yardstick

When schools create their own assessments, they have no way to ensure that one test is as challenging as another. A vocabulary test designed in one classroom may be significantly harder or easier than an identical test written by another teacher, not because of teaching quality, but simply due to how questions are phrased, selected, or weighted. This inconsistency makes it impossible to know whether a student's score reflects their actual learning or just the difficulty of the specific test they took.

The real damage happens when educators try to use these internally designed scores for comparison. You cannot reliably compare one student's performance to another's if both took different tests, even if those tests supposedly cover the same material. You cannot track genuine progress over time if assessment difficulty fluctuates from term to term. And you certainly cannot compare results fairly across different schools or subjects, where internal assessment standards vary widely.

How standardized tests create a level playing field

Standardized assessments work differently. They convert raw scores—the number of questions a student gets right—into a scaled score anchored to a standard distribution. This process, refined through rigorous statistical analysis, allows a student's performance to be meaningfully positioned against all other students who took the same test nationally.

This scaling mechanism unlocks genuine insight. A score of 75 on a standardized test tells you not just that a student answered 75% of questions correctly, but exactly where they sit relative to their peers across the entire country. It enables you to measure real growth over time—because the scale remains constant. And it allows fair comparison across different subjects and schools, because all scores are calibrated to the same standard. As Becky Sinjin explains in the episode, this consistency is what transforms raw data into actionable intelligence.

"The standardized test results are there to support the teacher judgment, to contribute to the teacher judgment, not to replace it."

Becky Sinjin — Freelance school assessment and performance data consultant. With 15 years as a data manager in secondary schools using CISRA tracking systems and seven years of freelance consulting work with multiple schools and companies on data management and assessment strategy, Sinjin brings deep operational expertise to the intersection of assessment design and data interpretation. She previously trained educators at GL Assessment on assessment systems before transitioning to independent practice.

This clarification—that standardized results support rather than replace professional judgment—highlights why the comparison gap matters. Teachers need reliable data to make better decisions faster, whether in selecting intervention students, tracking cohort progress, or identifying learning gaps. A standardized test provides that foundation; internally designed assessments cannot.

The broader implication, discussed in depth across this podcast episode, is that assessment design is not simply an operational detail—it directly shapes what educators can know about their students and, by extension, how effectively they can teach.

See also

What theoretical foundation supports the use of oracy in cognitive development?

Vygotsky's theories emphasize that communication tools are critical for cognitive development because they help shape the way individuals understand and process information.

How does oracy development support other areas of literacy like writing and reading?

Being able to talk through ideas before putting pen to paper is incredibly useful; 86% of staff at one secondary school believe oracy strategies supported overall literacy development.

How have teachers' initial reservations about oracy development changed over time in the trust?

There were initial teacher reservations when first starting to develop oracy in schools. As teachers learned and developed knowledge around oracy, speech and communication confidence increased significantly.

Key takeaways

Listen to the episode on Listenly