Open book with floating multiple-choice questions, symbolizing learning assessment

Are Your Exams Really Testing Knowledge? Uncover Hidden Flaws in Teacher-Made Tests

"Dive into a critical analysis of exam validity and cognitive assessment, revealing how to create fairer and more effective tests for academic success."


In Nigeria, there's growing concern about students not doing well on tests, especially big public ones. This makes people wonder if teachers are really preparing students well enough in the classroom. One important way teachers check how students are doing is through classroom tests. These tests give information that helps teachers know how well each student is learning. After the tests, the results are shared, and this feedback is helpful for students, their families, and the government, all of whom have a stake in education.

Classroom tests not only help to see where students stand but also let teachers know if they are likely to meet the goals for teaching a subject at a certain level. When these tests are done and results are out, everyone involved can look at how good their work was. Then, if needed, they can plan ways to help students improve or keep doing well.

One good thing about teacher-made tests is that teachers can pick their own words and set up the questions as they like, especially with multiple-choice questions. But because of this, there can be many different ways the questions are written, like how long the questions are, how many choices there are, and how clear the choices are. Also, some tests used in schools might not be checked well enough before they are used. So, it's important to look at how these tests are made and if they are really doing a good job of measuring what students know.

AI Search Multiple angles on this topic

The Case for Statistical Scrutiny

An analysis of teacher-made tests and teacher testing practices, as documented in an ERIC research report, led to the rejection of six stated null hypotheses, signaling significant quality concerns in how teachers construct and administer assessments. Free statistical calculators for hypothesis testing—including t-tests, ANOVA, and chi-square—are widely available to researchers in the social sciences, yet their use in evaluating teacher-made tests remains inconsistent. Statistical tools can reveal hidden flaws in test design, but many teachers never apply them to their own exams.

Standardized vs. Teacher-Made: A Tension in Testing

Standardized achievement tests provide norms and impartial information, while teacher-made tests are designed to evaluate teaching effectiveness but suffer from less accuracy and refinement. Both types of achievement tests are considered important for measuring student learning outcomes, yet they serve fundamentally different purposes. Research examining the validity of teacher-made tests for senior high school students highlights ongoing challenges in test construction, including ensuring content validity and alignment with curricular goals. Teachers in diverse educational contexts, such as Indonesia, face considerable obstacles when attempting to develop valid and reliable assessments.

Tracing the Roots of Assessment

The word 'history' itself derives from terms meaning inquiry and knowledge of past events, reflecting humanity's long effort to record and interpret significant occurrences. While the formalized study of educational testing emerged over centuries of scholarly effort, the specific historical milestones in the development of teacher-made tests as a field of inquiry remain under-documented in easily accessible sources. The broader trajectory of assessment—moving from informal evaluation to structured testing—parallels larger movements in education toward standardization and accountability.

Unveiling the Layers of Learning: Cognitive Domains and Test Design

Open book with floating multiple-choice questions, symbolizing learning assessment

Bloom's Taxonomy splits learning into six levels that show how deeply someone understands what's being taught. These levels include remembering facts, understanding them, using them, breaking things down, putting them together in new ways, and judging their value. The first three levels—remembering, understanding, and using—are seen as simpler thinking skills. Most tests in primary and junior secondary schools should focus on these because students at this level are still learning how to think critically and abstractly. It wouldn't be right to ask them to do tasks that need more advanced thinking.

A well-made multiple-choice question (MCQ) should test students on all six areas of thinking, as Bloom suggested. If a test doesn't cover all these areas, it might not truly show what a student knows. A good test should have a mix of questions from different levels, with fewer questions on the very basic and very advanced levels and more in the middle. To do this well, teachers can use a Test Blueprint, which helps them plan the test. However, many teachers don't use this blueprint, even though it's a key to making sure the test is fair and covers everything. Even if teachers are busy, using a test blueprint is a must for creating good tests that check all levels of thinking.

Key Objectives of This Analysis:
  • Determine if the selected undergraduate tests accurately measure course content.
  • Examine the distribution of test questions across Bloom's Taxonomy levels.
  • Assess if the answer choices in the tests are reasonable and well-constructed.
AI Search Multiple angles on this topic

What Recent Studies Reveal About Teacher-Made Tests

Research published in the South African Journal of Education finds that teacher-made tests are highly vulnerable to misuse by teachers who adopt a haphazard approach to test planning, administration, and scoring. Studies of EFL teacher-made tests confirm that teachers frequently default to multiple-choice formats when developing assessments, often without rigorous item analysis. In Nkayi District primary schools, researchers found that teacher-made tests serve the important function of helping teachers identify content mastered by pupils and areas where students struggle. However, the negative backwash effect of both centralized and teacher-made tests has been shown to downplay teachers' professional agency in favor of dominant institutional structures.

Pushback Against the Testing Paradigm

Critics of the Zone of Proximal Development argue that even if the concept is theoretically sound, it is extremely difficult to measure in practice—a challenge that extends to any assessment framework, including teacher-made tests. Political figures have also entered the debate, with Hillary Clinton publicly criticizing the overuse and abuse of standardized tests in an effort to gain support from teachers' unions. National Education Association leaders have spoken extensively about how excessive testing undermines teaching quality, a concern that applies equally to teacher-made and standardized exams. The broader argument against testing culture raises questions about whether any single assessment format can truly capture student learning.

Traditional vs. Alternative Assessment: Incompatible Formats?

A study documented in an ERIC report compared traditional and alternative reading and math scores, supplemented by surveys of 20 teachers and 100 students. The results indicated that traditional and alternative types of testing cannot be compared the majority of the time, pointing to a fundamental incompatibility between assessment methods. This finding suggests that teacher-made tests (a form of traditional assessment) and alternative assessments measure different constructs and should not be used interchangeably. The research underscores a clear need for both types of assessments to be employed, as neither alone can provide a complete picture of student achievement.

The study looked at tests from Obafemi Awolowo University in Nigeria. It included tests from different courses that many students take. Researchers looked at how well the tests covered the material and if the questions were appropriate for the students' learning levels. They also checked if the answer choices were clear and made sense. This helped them understand if the tests were really measuring what the students knew.

Key Recommendations for Universities

Universities and similar institutions should create tests that align with the topics listed in the course content. It's also important to have more questions per course, aiming for between 70 and 100. This improves the test's reliability and better reflects the course's objectives. Additionally, including more questions that assess high-level thinking skills ensures that students who pass the courses can demonstrate a high level of proficiency.

AI Search Multiple angles on this topic

Expert Perspectives on Teacher-Made Testing

James B. Carter has argued that teacher-made tests remain among the best tools available for understanding students, noting that teachers across Rowan-Salisbury Schools use ice-breakers, questionnaires, journal entries, and assessments to reveal students' strengths and areas for improvement. The hands-on, relationship-driven approach to assessment contrasts sharply with the impersonal nature of large-scale standardized testing. Carter's perspective suggests that the value of teacher-made tests lies not just in measurement but in the deeper teacher-student connection that informs test design and interpretation. Expert opinion thus favors teacher-made tests when they are used thoughtfully and as part of a holistic understanding of each learner.

Digital Tools and the Evolution of Test Creation

Digital platforms such as TeacherMade are emerging to streamline the process of creating, administering, and grading teacher-made assessments online. These tools promise to reduce the haphazard approach that research has identified as a key weakness of teacher-made tests by providing structured templates and automated scoring features. As educational technology continues to evolve, the gap between the rigor of standardized tests and the convenience of teacher-made tests may narrow. The future of teacher-made testing will likely depend on whether educators embrace such tools and integrate them with sound assessment principles.

Teacher Burnout and the Testing Burden

Research on teacher burnout identifies excessive workload as a key stressor, alongside lack of administrative support and systemic pressures that affect classroom management and teacher-student relationships. The burden of designing, administering, and scoring tests—whether teacher-made or standardized—contributes to this workload crisis. Institutional cohesion suffers when teachers are overwhelmed by assessment demands that pull them away from instructional priorities. Addressing the hidden flaws in teacher-made tests requires acknowledging the systemic conditions under which those tests are produced.

How Test Origin Shapes Student Outcomes

A case study comparing the quality of teacher-made tests with achievement tests found meaningful discrepancies in student scoring depending on whether the test was created by the educator or by an external researcher. These differences highlight how the identity and approach of the test designer can directly influence pupils' measured success. Students may perform differently on teacher-made tests due to familiarity with the teacher's style, question framing, or content emphasis. The findings underscore that test quality is not merely a technical issue but one with real consequences for how student achievement is perceived and reported.

About this Article -

Written with AI assistance from published research, and reviewed by the Mystum team. See our About page for more information.

Everything You Need To Know

1

How do teacher-made tests contribute to student assessment and educational feedback?

Teacher-made tests are used to check student progress. These tests provide insights into individual student learning, which is helpful feedback for students, families, and the government. They assess whether students are likely to meet course objectives. The feedback helps stakeholders evaluate the effectiveness of their work and plan for improvements.

2

Can you explain Bloom's Taxonomy and its role in designing effective tests?

Bloom's Taxonomy divides learning into six levels: remembering, understanding, applying, analyzing, evaluating, and creating. The initial three levels (remembering, understanding, and applying) represent simpler cognitive skills, whereas the latter three involve higher-order thinking. Assessments should include a mix of questions across these levels to comprehensively evaluate student knowledge, with a focus on the lower levels for primary and junior secondary students.

3

What is a Test Blueprint, and why is it important for creating fair and comprehensive tests?

A Test Blueprint is used to plan and structure tests. It ensures fair coverage of all learning objectives and cognitive levels outlined in Bloom's Taxonomy. By using a Test Blueprint, educators can create balanced assessments that accurately measure what students know. Even with time constraints, the use of a Test Blueprint is essential for test quality.

4

What parameters made Obafemi Awolowo University in Nigeria an ideal case study for test analysis?

Obafemi Awolowo University in Nigeria was selected as a case study. Tests from various commonly taken courses were examined to assess content coverage, appropriateness for students' learning levels, and the clarity of answer choices. The goal was to determine if these tests accurately measure students' knowledge, and identify any shortcomings in test design and implementation.

5

What recommendations are provided to universities to improve their test design and ensure that students demonstrate high-level proficiency?

For universities, tests should align with course content and include a higher number of questions, ideally between 70 and 100 per course, to enhance reliability. Furthermore, there should be a greater emphasis on questions that assess higher-level thinking skills, ensuring that successful students demonstrate proficiency in advanced cognitive processes. This helps universities make sure that they are measuring the course content correctly.

Newsletter Subscribe

Subscribe to get the latest articles and insights directly in your inbox.