Are Your Exams Really Testing Knowledge? Uncover Hidden Flaws in Teacher-Made Tests
"Dive into a critical analysis of exam validity and cognitive assessment, revealing how to create fairer and more effective tests for academic success."
In Nigeria, there's growing concern about students not doing well on tests, especially big public ones. This makes people wonder if teachers are really preparing students well enough in the classroom. One important way teachers check how students are doing is through classroom tests. These tests give information that helps teachers know how well each student is learning. After the tests, the results are shared, and this feedback is helpful for students, their families, and the government, all of whom have a stake in education.
Classroom tests not only help to see where students stand but also let teachers know if they are likely to meet the goals for teaching a subject at a certain level. When these tests are done and results are out, everyone involved can look at how good their work was. Then, if needed, they can plan ways to help students improve or keep doing well.
One good thing about teacher-made tests is that teachers can pick their own words and set up the questions as they like, especially with multiple-choice questions. But because of this, there can be many different ways the questions are written, like how long the questions are, how many choices there are, and how clear the choices are. Also, some tests used in schools might not be checked well enough before they are used. So, it's important to look at how these tests are made and if they are really doing a good job of measuring what students know.
The Case for Statistical Scrutiny
An analysis of teacher-made tests and teacher testing practices, as documented in an ERIC research report, led to the rejection of six stated null hypotheses, signaling significant quality concerns in how teachers construct and administer assessments. Free statistical calculators for hypothesis testing—including t-tests, ANOVA, and chi-square—are widely available to researchers in the social sciences, yet their use in evaluating teacher-made tests remains inconsistent. Statistical tools can reveal hidden flaws in test design, but many teachers never apply them to their own exams.
Standardized vs. Teacher-Made: A Tension in Testing
Standardized achievement tests provide norms and impartial information, while teacher-made tests are designed to evaluate teaching effectiveness but suffer from less accuracy and refinement. Both types of achievement tests are considered important for measuring student learning outcomes, yet they serve fundamentally different purposes. Research examining the validity of teacher-made tests for senior high school students highlights ongoing challenges in test construction, including ensuring content validity and alignment with curricular goals. Teachers in diverse educational contexts, such as Indonesia, face considerable obstacles when attempting to develop valid and reliable assessments.
Tracing the Roots of Assessment
The word 'history' itself derives from terms meaning inquiry and knowledge of past events, reflecting humanity's long effort to record and interpret significant occurrences. While the formalized study of educational testing emerged over centuries of scholarly effort, the specific historical milestones in the development of teacher-made tests as a field of inquiry remain under-documented in easily accessible sources. The broader trajectory of assessment—moving from informal evaluation to structured testing—parallels larger movements in education toward standardization and accountability.
Unveiling the Layers of Learning: Cognitive Domains and Test Design
Bloom's Taxonomy splits learning into six levels that show how deeply someone understands what's being taught. These levels include remembering facts, understanding them, using them, breaking things down, putting them together in new ways, and judging their value. The first three levels—remembering, understanding, and using—are seen as simpler thinking skills. Most tests in primary and junior secondary schools should focus on these because students at this level are still learning how to think critically and abstractly. It wouldn't be right to ask them to do tasks that need more advanced thinking.
- Determine if the selected undergraduate tests accurately measure course content.
- Examine the distribution of test questions across Bloom's Taxonomy levels.
- Assess if the answer choices in the tests are reasonable and well-constructed.
What Recent Studies Reveal About Teacher-Made Tests
Research published in the South African Journal of Education finds that teacher-made tests are highly vulnerable to misuse by teachers who adopt a haphazard approach to test planning, administration, and scoring. Studies of EFL teacher-made tests confirm that teachers frequently default to multiple-choice formats when developing assessments, often without rigorous item analysis. In Nkayi District primary schools, researchers found that teacher-made tests serve the important function of helping teachers identify content mastered by pupils and areas where students struggle. However, the negative backwash effect of both centralized and teacher-made tests has been shown to downplay teachers' professional agency in favor of dominant institutional structures.
Pushback Against the Testing Paradigm
Critics of the Zone of Proximal Development argue that even if the concept is theoretically sound, it is extremely difficult to measure in practice—a challenge that extends to any assessment framework, including teacher-made tests. Political figures have also entered the debate, with Hillary Clinton publicly criticizing the overuse and abuse of standardized tests in an effort to gain support from teachers' unions. National Education Association leaders have spoken extensively about how excessive testing undermines teaching quality, a concern that applies equally to teacher-made and standardized exams. The broader argument against testing culture raises questions about whether any single assessment format can truly capture student learning.
Traditional vs. Alternative Assessment: Incompatible Formats?
A study documented in an ERIC report compared traditional and alternative reading and math scores, supplemented by surveys of 20 teachers and 100 students. The results indicated that traditional and alternative types of testing cannot be compared the majority of the time, pointing to a fundamental incompatibility between assessment methods. This finding suggests that teacher-made tests (a form of traditional assessment) and alternative assessments measure different constructs and should not be used interchangeably. The research underscores a clear need for both types of assessments to be employed, as neither alone can provide a complete picture of student achievement.
Key Recommendations for Universities
Universities and similar institutions should create tests that align with the topics listed in the course content. It's also important to have more questions per course, aiming for between 70 and 100. This improves the test's reliability and better reflects the course's objectives. Additionally, including more questions that assess high-level thinking skills ensures that students who pass the courses can demonstrate a high level of proficiency.
Expert Perspectives on Teacher-Made Testing
James B. Carter has argued that teacher-made tests remain among the best tools available for understanding students, noting that teachers across Rowan-Salisbury Schools use ice-breakers, questionnaires, journal entries, and assessments to reveal students' strengths and areas for improvement. The hands-on, relationship-driven approach to assessment contrasts sharply with the impersonal nature of large-scale standardized testing. Carter's perspective suggests that the value of teacher-made tests lies not just in measurement but in the deeper teacher-student connection that informs test design and interpretation. Expert opinion thus favors teacher-made tests when they are used thoughtfully and as part of a holistic understanding of each learner.
Digital Tools and the Evolution of Test Creation
Digital platforms such as TeacherMade are emerging to streamline the process of creating, administering, and grading teacher-made assessments online. These tools promise to reduce the haphazard approach that research has identified as a key weakness of teacher-made tests by providing structured templates and automated scoring features. As educational technology continues to evolve, the gap between the rigor of standardized tests and the convenience of teacher-made tests may narrow. The future of teacher-made testing will likely depend on whether educators embrace such tools and integrate them with sound assessment principles.
Teacher Burnout and the Testing Burden
Research on teacher burnout identifies excessive workload as a key stressor, alongside lack of administrative support and systemic pressures that affect classroom management and teacher-student relationships. The burden of designing, administering, and scoring tests—whether teacher-made or standardized—contributes to this workload crisis. Institutional cohesion suffers when teachers are overwhelmed by assessment demands that pull them away from instructional priorities. Addressing the hidden flaws in teacher-made tests requires acknowledging the systemic conditions under which those tests are produced.
How Test Origin Shapes Student Outcomes
A case study comparing the quality of teacher-made tests with achievement tests found meaningful discrepancies in student scoring depending on whether the test was created by the educator or by an external researcher. These differences highlight how the identity and approach of the test designer can directly influence pupils' measured success. Students may perform differently on teacher-made tests due to familiarity with the teacher's style, question framing, or content emphasis. The findings underscore that test quality is not merely a technical issue but one with real consequences for how student achievement is perceived and reported.