Teacher evaluation systems answer a difficult measurement problem: how can a school or education authority judge teaching quality without reducing it to one lesson, one test score, or one supervisor’s opinion? Countries reach very different answers. Some rely on annual meetings led by the principal, some use national teaching standards, and others connect portfolios and subject-knowledge tests to career progression. The design matters because an appraisal can serve two distinct purposes: improving practice through feedback and making formal decisions about recognition, advancement, support, or employment.
What International Data Shows
Published in October 2025, the OECD’s 2024 Teaching and Learning International Survey is the newest large cross-country dataset available in 2026. It collected responses from 280,000 educators in 55 education systems. On average, 88% of lower-secondary teachers worked in schools where the principal formally appraised teachers at least once a year. Classroom observation was used in schools covering 96% of teachers. Yet the follow-up was less consistent: 65% of teachers worked in schools where appraisal led to a discussion about improving teaching, 46% had access to a development or training plan, and only 20% could receive mentoring. Financial incentives reached 12%, while formal sanctions affected fewer than 3%.[a]
Teacher Appraisal Is Not the Same as School Inspection
Teacher appraisal assesses an individual educator’s professional practice. It may examine lesson planning, instruction, assessment literacy, classroom climate, professional learning, collaboration, and wider responsibilities. The evaluator is usually a principal, department head, mentor, trained peer, external assessor, or a combination of these people.
School evaluation or inspection examines the institution. It considers leadership, curriculum delivery, safeguarding, inclusion, student outcomes, governance, and organizational quality. Evidence about teaching may inform a school-level judgment, but that does not automatically produce a rating for each teacher. Japan, for example, has legal provisions for school self-evaluation, publication of results, stakeholder evaluation, and reporting to the school’s establishing authority; those processes operate alongside personnel evaluation rather than replacing it.[b]
Student assessment is different again. Tests, course grades, progress measures, and classroom work describe student learning. They can contribute evidence about instruction, but they also reflect prior attainment, attendance, language background, family resources, curriculum alignment, peer effects, and access to specialist support. A credible evaluation system treats student results as contextual evidence rather than a complete measurement of teacher quality.
How National and Local Models Differ
| Education System | Main Level of Control | Typical Evidence | Primary Use |
|---|---|---|---|
| Finland | Municipality and school | Professional dialogue, local goals, principal feedback | Development and local management |
| England | Statutory national rules with school policies | Objectives, professional standards, observation, other agreed evidence | Annual appraisal and professional development |
| Australia | States, territories, sectors, and schools within a national standards model | Goals, observations, student learning evidence, feedback, professional learning | Performance and development |
| Singapore | Ministry and school leadership | Work performance, professional contribution, feedback, career readiness | Development, recognition, and career decisions |
| Japan | Local boards of education and school leadership | Goals, duties, professional competence, school priorities | Personnel management and improvement |
| Chile | National career system | Portfolio, recorded lesson, collaborative work, subject and pedagogical knowledge | Career-stage recognition and progression |
| United States | State and local district | Observation rubrics, professional goals, local student measures, other evidence | Feedback, employment, and advancement under local rules |
Finland: Local Responsibility and Professional Autonomy
Finland does not center teacher quality assurance on a single national rating instrument. Municipal employers and school leaders manage staff performance locally, while national curriculum requirements, teacher qualifications, school self-evaluation, and sample-based system evaluation provide the wider quality structure. This model gives principals and municipalities room to shape feedback around school needs.
TALIS 2024 makes the lower frequency visible. Among Finnish teachers reporting substantial or full autonomy in curriculum implementation, 53% worked in schools where they were appraised less than annually or not at all. Across all Finnish lower-secondary teachers, 62% worked in schools where the principal conducted formal appraisal at least annually, compared with an OECD average of 88%. Fewer than 10% worked in schools reporting formal peer appraisal. These figures describe frequency, not an absence of professional responsibility: qualification requirements, collegial planning, principal oversight, curriculum duties, and local employment processes still shape practice.[c]
The Finnish case shows why appraisal intensity and system quality are not interchangeable. A lower-frequency model places more weight on professional preparation, trust, and routine school leadership. Its measurement challenge is consistency: teachers may receive very different feedback depending on the capacity and habits of the municipality or principal.
England: Annual Appraisal under Statutory Rules
England uses a more explicit annual process for teachers in local-authority-maintained schools. The statutory setting is national, while schools develop appraisal policies and conduct the process. Teachers normally receive objectives, have their performance reviewed against their role and relevant professional standards, and receive a written appraisal report. Classroom observation may contribute evidence, but it should sit within a broader body of information rather than act as a stand-alone verdict.
Department for Education guidance updated for September 2024 separates routine appraisal from capability procedures addressing serious underperformance. That separation is important. Developmental appraisal works best when ordinary improvement needs do not automatically trigger a disciplinary process. The guidance applies to maintained schools and local authorities in relation to unattached teachers; academies operate with greater policy discretion, although many use comparable annual cycles.[d]
Australia: Standards-Based Performance and Development
Australia combines decentralized employment arrangements with a shared professional reference point. The Australian Professional Standards for Teachers describe practice across four career stages: Graduate, Proficient, Highly Accomplished, and Lead. The national performance and development model, endorsed by education ministers in 2012, expects a recurring cycle of goal setting, professional practice and learning, feedback, and review.
The national model does not create one score for every teacher. States, territories, Catholic and independent sectors, employers, and schools implement processes within their own arrangements. Evidence may include observation, student work and learning data, planning documents, feedback from colleagues, and records of professional learning. Its central design feature is alignment between goals, evidence, feedback, and targeted development, with the standards supplying a shared vocabulary.[e]
Singapore: Appraisal Connected to Career Development
Singapore’s Ministry of Education integrates appraisal with a centrally managed teaching career. School leaders and appointed personnel review performance, give feedback, identify support needs, and consider readiness for greater responsibilities. The process therefore has both developmental and personnel functions.
Ministry statements confirm that performance appraisal informs feedback and support, while consistently good performance and readiness for a higher-level role can inform promotion. Teachers who need improvement continue to receive supervisor guidance and relevant in-service training. The ministry also states that 360-degree feedback is used mainly as a developmental tool for leaders rather than as a direct appraisal instrument. Singapore retains relative ranking guidelines but allows contextual departures and periodically reviews the model.[f]
This arrangement makes career consequences more visible than in low-stakes local models. It also requires careful calibration among supervisors. When appraisal influences promotion, evaluators must distinguish current performance, leadership potential, assigned opportunities, and the visibility of different kinds of work.
Japan: Evaluation through Local Education Authorities
Japan’s public-school teachers are employed within a system where prefectural and designated-city boards of education hold major personnel responsibilities. Local systems commonly combine teacher-set goals with assessment by school leaders. Areas considered can include instructional work, student guidance, school management duties, professional competence, and contribution to school objectives.
The Ministry of Education, Culture, Sports, Science and Technology has documented the rollout of new teacher evaluation systems across local boards. In an earlier national implementation snapshot, 57 of 62 prefectural and designated-city boards were piloting or operating new systems by April 2006. The historical figure marks a broad institutional shift, not a current annual participation rate. Present procedures still vary locally, so Japan is better understood as a nationally guided but locally administered personnel model than as one uniform evaluation form.[g]
Chile: Portfolios and Knowledge Assessment for Career Progression
Chile offers one of the clearest examples of nationally structured, evidence-rich evaluation. Its Recognition and Promotion System, created under Law 20.903 and updated through Law 21.625 of 2023, links recognized career stages to professional experience, the teacher’s existing stage, a pedagogical portfolio, and the Evaluation of Specific and Pedagogical Knowledge. Career stages include Initial, Early, Advanced, Expert I, and Expert II. Since 2024, the unified recognition process has replaced the former double-evaluation arrangement.[h]
The portfolio gathers direct evidence of practice through three broad modules: planning, assessment and reflection; a recorded class; and collaborative work. Trained educators score the evidence against published rubrics. Performance indicators use four descriptive levels, while the final portfolio result is translated into categories used with the knowledge assessment and experience record for career placement. The knowledge assessment contains 60 multiple-choice questions covering disciplinary and pedagogical knowledge for the relevant subject, level, or specialization.[i]
Chile’s approach captures evidence that a short school visit cannot. It can examine planning decisions, classroom interaction, reflection, collaboration, and content knowledge. The cost is time: portfolio production, scorer preparation, quality control, appeals, platform administration, and test development create a larger operational load than a principal-led annual conversation.
United States: State and District Variation
The United States has no single national teacher evaluation system. State law sets many parameters, but districts often select rubrics, observation schedules, professional goals, student-growth measures, and improvement procedures. Requirements may also differ for probationary and experienced teachers. A teacher moving between states can therefore encounter a materially different process.
Student Learning Objectives became one method for incorporating learning growth where standardized test-based measures were unavailable or unsuitable. A federal review identified SLO use in 30 states in 2014, illustrating the scale of that policy period without implying that every system retained the same design. Current local models often combine observations and professional goals with multiple evidence sources, but weighting and consequences remain state- and district-specific.[j]
TALIS 2024 provides a newer cross-section. In participating U.S. schools, principals reported classroom observation as an appraisal method for 100% of teachers, school- or classroom-based results for 93%, and external student results for 89%. Following appraisal, 81% worked in schools where weaknesses could prompt a discussion, 54% where a development plan could be created, and 30% where a mentor could be appointed.[k]
Evidence Used to Judge Teaching Quality
Classroom Observation
Observation supplies direct evidence of instruction: clarity of explanation, questioning, checks for understanding, classroom climate, subject accuracy, pacing, and adaptation. It is the most widely used appraisal method internationally. Its weakness is sampling. One lesson represents a small portion of a teacher’s year, and the observer, class composition, subject, lesson phase, and whether the visit was announced can affect the rating.
- Reliability improves with trained observers, shared rubrics, calibration exercises, and more than one observation.
- Validity improves when the rubric describes observable teaching rather than vague personality traits.
- Usefulness improves when feedback names evidence, explains its effect on learning, and identifies a feasible next step.
Portfolios and Professional Artifacts
Portfolios may include lesson sequences, assessments, samples of student work, feedback records, reflective commentary, collaboration evidence, and video. They reveal decisions made before and after a lesson, which observation alone cannot capture. They also create comparability problems when teachers have unequal time, technical assistance, class assignments, or experience in presenting evidence. Clear rubrics and authenticity controls are therefore essential.
Student Learning Evidence
Student evidence ranges from classroom work and curriculum-based assessments to external examinations and statistical growth estimates. The closer the measure is to the taught curriculum, the easier it is to interpret instructionally. Yet narrow measures can encourage excessive focus on tested content, while broad results may be too distant from one teacher’s contribution.
Value-added models attempt to estimate a teacher’s contribution after accounting for prior achievement and selected student characteristics. Their precision depends on data quality, model specification, sample size, student-teacher linkage, test scaling, and the stability of estimates across years. They also cover only tested grades and subjects. For these reasons, a numerical growth estimate should be interpreted with other evidence and an uncertainty interval, not treated as a perfectly exact rank.
Student and Parent Surveys
Students observe many lessons and can report whether explanations are clear, expectations are consistent, feedback is useful, and the classroom supports participation. Surveys can therefore add a perspective that a visiting evaluator cannot. They should use age-appropriate, tested questions and minimum response thresholds. Popularity, grading strictness, response bias, and unequal expectations may distort individual scores, so surveys are better suited to patterns than isolated comments.
Self-Assessment, Peer Review, and Professional Learning
Self-assessment helps teachers connect evidence to goals, while peer review brings subject and grade-level expertise into the process. Records of professional learning show participation, but attendance alone does not demonstrate changed practice. Evaluators need evidence of application: revised planning, new assessment methods, stronger student work, or a documented improvement in classroom routines.
The Technical Tests of a Fair Evaluation System
| Criterion | Question It Answers | Common Design Response |
|---|---|---|
| Validity | Does the process measure teaching practice relevant to the role? | Align evidence with professional standards and actual duties. |
| Reliability | Would trained evaluators reach reasonably consistent judgments? | Use calibration, multiple observations, and scoring checks. |
| Comparability | Are ratings interpreted similarly across schools and subjects? | Use shared descriptors, exemplars, and moderation. |
| Context Sensitivity | Does the judgment account for the teacher’s assignment and available support? | Record class composition, resources, timetable, and role scope. |
| Actionability | Can the teacher act on the feedback? | Connect findings to time, coaching, mentoring, or training. |
| Procedural Fairness | Can the teacher understand and respond to the decision? | Publish criteria, disclose evidence, and provide review routes. |
Ratings Need More than a Rubric
A detailed rubric does not guarantee consistent scoring. Evaluators can differ in severity, favor certain teaching styles, or place too much weight on the most recent event. Halo effects occur when one strong characteristic lifts every domain; confirmation bias appears when an early impression shapes later evidence. Calibration sessions address these risks by asking evaluators to score common lesson videos or artifacts, compare reasoning, and resolve differences against the rubric.
Moderation also matters after scoring begins. Systems can examine rating distributions by school, evaluator, subject, career stage, and teacher group. A difference does not prove bias, but a persistent unexplained pattern signals the need for review. Inter-rater agreement, score stability, missing-evidence rates, appeals, and completion times are useful technical indicators.
Context Must Be Recorded without Lowering Expectations
Teachers work under different conditions. Class size, subject, student age, prior attainment, language needs, special education support, staff vacancies, timetable fragmentation, and access to materials can change what observers see and what student measures report. Context should explain the evidence, not erase professional standards. The distinction is similar to reading a thermometer with its location noted: the number remains real, but interpretation depends on where and how it was measured.
Fairness also requires role-specific evidence. A special education teacher, vocational instructor, counselor-teacher, early-years educator, and upper-secondary subject specialist do not produce identical artifacts. Shared professional principles can remain constant while indicators and evidence routes vary.
What Happens after the Rating Matters More than the Form
Appraisal creates value only when a school can respond. A teacher may need coaching in questioning, time to observe an experienced colleague, subject-specific training, help analyzing student work, or a reduced set of priorities. A generic recommendation to “improve instruction” offers no usable direction.
The TALIS 2024 gap between annual appraisal and subsequent support is therefore revealing. Formal review reaches most teachers, but fewer than half work in schools offering a development or training plan after appraisal. In Denmark, Finland, France, Japan, Korea, and Portugal, fewer than half worked in schools reporting post-appraisal discussions aimed at addressing teaching weaknesses. The measure does not show the quality of informal feedback, but it does identify a weak connection between judgment and organized support in many settings.
- Evidence collection identifies a specific aspect of practice.
- Professional dialogue tests the interpretation with the teacher.
- A focused goal defines the expected change.
- Support supplies time, expertise, materials, or coaching.
- Follow-up evidence shows whether practice changed.
Systems also need a boundary between ordinary development and formal capability action. Every teacher has areas to refine; treating any development goal as evidence of failure discourages honest reflection. Formal intervention is more defensible when expectations, evidence, support already offered, timelines, decision authority, and review rights are clearly documented.
Teacher Evaluation, Workload, and Trust
Evaluation carries an administrative cost. Teachers prepare evidence, observers leave other duties, leaders write reports, and systems maintain records. The burden grows when schools collect artifacts that no one uses or duplicate information already available. A lean system asks for each item only when it can change feedback, support, or a formal decision.
Accountability pressure can also affect well-being. Across OECD education systems in TALIS 2024, 45% of teachers reported that responsibility for student achievement caused “quite a bit” or “a lot” of stress. The share was below one-third in Finland, Hungary, Iceland, and Kazakhstan but above 70% in Latvia, Lithuania, Portugal, and South Africa. These are perceptions rather than causal estimates, yet they show that the same language of accountability can be experienced very differently across systems.
Trust is not the absence of evidence. It is confidence that evidence will be relevant, interpreted competently, and used according to stated rules. High-trust evaluation still needs standards and documentation. High-accountability evaluation still needs professional voice, context, and meaningful support.
Digital Records and AI Create New Measurement Questions
Digital portfolios, video observation, learning platforms, and automated dashboards make it easier to collect evidence at scale. They also increase the amount of personal data held about teachers and students. Systems need defined retention periods, access controls, correction procedures, and rules governing who can use recordings or analytics. A classroom video created for appraisal should not silently become material for unrelated purposes.
Generative AI adds a newer problem. It can help organize notes or summarize documents, but an opaque model should not determine a teacher’s rating. Automated outputs may misread subject-specific practice, reproduce patterns in historical ratings, or infer quality from language style rather than instruction. Human evaluators must remain accountable for the judgment, and teachers should know when an automated system has processed their evidence.
Before any digital measure enters a high-stakes process, authorities need evidence of accuracy, consistency across groups and subjects, data security, and practical relevance. A convenient metric is not automatically a valid metric.
How the Main Models Compare
Country systems can be grouped by the decision they are designed to support. The categories overlap, but they clarify why countries select different evidence and schedules.
- Professional-growth models emphasize feedback, goals, coaching, and local judgment. Finland reflects the lower-frequency end of this family.
- Standards-based annual models use a recurring school-led cycle tied to national or jurisdictional teaching standards. England and Australia illustrate different versions.
- Career-management models connect appraisal to promotion, role readiness, or differentiated career pathways. Singapore places appraisal within a centralized career structure.
- National certification or recognition models use common instruments and external scoring to support career-stage decisions. Chile’s portfolio and knowledge assessment provide a clear example.
- Decentralized accountability models allow state, district, or local authorities to determine instruments and weights. The United States shows the widest internal variation among the examples examined.
No model removes trade-offs. Local discretion improves contextual fit but can weaken comparability. National instruments improve consistency but add workload and may fit some subjects better than others. Developmental systems encourage candor but may offer limited information for promotion decisions. High-stakes systems clarify consequences but can encourage strategic presentation of evidence.
What a Credible System Ultimately Measures
A credible evaluation system does not search for one universal indicator. It combines direct evidence of teaching, professional knowledge, student learning information, reflection, and contribution to the school in proportions suited to the decision being made. The evidence needed for a coaching conversation is not identical to the evidence needed for national career-stage recognition.
The strongest common principle across countries is alignment. Standards must match the teacher’s role; instruments must match the standards; evaluator training must match the instruments; consequences must match the certainty of the evidence; and professional support must match the diagnosed need. Where one link fails, evaluation becomes paperwork or a rating exercise. Where the links hold, it becomes a disciplined way to strengthen teaching while making personnel decisions more transparent.
Sources
- [a] Results from TALIS 2024: The State of Teaching: OECD data on appraisal frequency, methods, follow-up support, incentives, sanctions, and accountability-related stress.
- [b] 学校評価に関する学校教育法・学校教育法施行規則の規定: MEXT provisions on school self-evaluation, stakeholder evaluation, publication, and reporting.
- [c] Results from TALIS 2024 – Country Notes: Finland: Finland’s appraisal frequency, autonomy, professional relationships, and feedback indicators.
- [d] Teacher Appraisal and Capability: Department for Education guidance for appraisal and capability policies in England.
- [e] AITSL Performance and Development Resource: Australia’s national structure for goals, feedback, professional learning, and review.
- [f] Performance Appraisal Models for Teachers and Teachers’ Performance Management: Singapore Ministry of Education statements on ranking, feedback, promotion, support, and leadership feedback.
- [g] 教員評価について: MEXT information on the implementation of teacher evaluation systems by local education boards.
- [h] Sistema de Reconocimiento: Chile’s official description of career stages, legal basis, portfolio, knowledge assessment, and the unified evaluation process.
- [i] Portafolio and Evaluación de Conocimientos Específicos y Pedagógicos: official details on Chile’s portfolio modules, rubrics, scoring, feedback, and 60-question knowledge assessment.
- [j] How States Use Student Learning Objectives in Teacher Evaluation Systems: A Review of State Websites: U.S. Institute of Education Sciences review of SLO use in state evaluation systems.
- [k] Results from TALIS 2024 – Country Notes: United States: current indicators for U.S. appraisal methods and post-appraisal responses.