Conference Presentation, Panel, Fireside Chat
Teacher Evaluation and National Testing: Can We Reach Consensus?
Milken InstituteYvonne Chan, Jason Culbertson, Warren Fletcher, Patrice Pujol, Thomas Boysen, Tom Boyce
Context and Strategic Shift
- The panel discussed the transition from viewing social class as the primary determinant of student success (1966 Coleman Report) to recognizing teacher effectiveness as the critical variable.
- Research indicates that two years with highly effective teachers can move a student from the 50th to the 90th percentile, whereas two years with ineffective teachers can drop them to the 37th percentile.
- The Gates Foundation's $45 million, multi-year study concluded that teaching can be effectively measured using three specific data points: test score growth, student feedback, and multiple observations against a clear rubric.
- Economic analysis suggests a highly effective teacher generates approximately $250,000 in lifetime income value per classroom.
- Current teacher compensation models heavily reward years of service and degree attainment, neither of which strongly correlates with teacher effectiveness or student outcomes.
- High teacher turnover is identified as a significant financial and operational cost to the education system.
Policy and Standardization Trends
- The implementation of Common Core State Standards and associated assessments (scheduled for 2014-2015) aims to resolve "inflation-adjusted achievement" where state tests showed higher results than the National Assessment of Educational Progress (NAEP).
- Political consensus on national standards has been difficult, with historical resistance from both political parties regarding the words "national" and "testing."
- The MET (Measuring Effective Teaching) study data influenced the requirement for multiple observations rather than single, subjective walkthroughs.
Case Study: Vaughan Learning Center (Yvonne Chan)
- Challenge: In 1997, the school faced a 60% annual teacher attrition rate and a 50% workforce without teaching credentials in a high-poverty, high-English-learner environment (85% EL, 100% Title I).
- Solution: Implemented a comprehensive, standard-based evaluation system involving 20 observable dimensions, clear rubrics, and peer collaboration across preschool through high school.
- Incentive Structure: Introduced performance pay up to $15,000 annually for teachers and administrators, linked to professional growth rather than just compliance.
- Cultural Shift: The system requires collaboration across disciplines (e.g., preschool teachers responsible for high school literacy outcomes) to ensure school-wide success.
- Scalability: The model successfully phased in new hires first, followed by veterans and administrators, allowing for iterative refinement of the rubric every two to three years.
Case Study: United Teachers of Los Angeles (Warren Fletcher)
- Legal Mandate: A court ruling (Doe v. Dacey) found LAUSD inconsistently used test data, ordering the union and district to collectively bargain a new system incorporating student achievement data within six months.
- Bargaining Philosophy: The union rejected a rigid Value-Added Model (VAM) that simply assigned a score, arguing it offered no roadmap for improvement.
- Preferred Model: Advocated for a system where data identifies specific instructional gaps (e.g., difficulty teaching "graphing" vs. "factoring") to enable targeted coaching and resource allocation.
- Quality Control: Emphasized the necessity of calibrated evaluators to prevent bias, noting that uncalibrated "360-degree" evaluations can be manipulated or subjective.
- Risk Warning: Fiercely opposed tying pay solely to test scores, warning it would incentivize "teaching to the test" (e.g., drilling 5-paragraph essay structures) at the expense of deep literary analysis and complex writing.
- Union Position: Argued that unions must focus on professional practice and student outcomes, not just work rules, to maintain credibility and improve the profession.
Case Study: Ascension Parish Schools (Patrice Pujol)
- Objective: Utilize teacher effectiveness systems to close the achievement gap between affluent and impoverished schools within the district.
- Implementation Strategy: Expanded a pilot turnaround zone (8 schools) using the TAP system to all 28 district schools.
- Evaluators: Trained 500 of 1,400 teachers as evaluators to conduct peer observations, ensuring feedback comes from instructional experts rather than solely administrators.
- Feedback Loop: Shifted from end-of-year summative evaluations to weekly, monthly, and quarterly tracking of student outcomes to provide immediate, actionable feedback.
- Cultural Integration: Relied on professional pressure and peer collaboration to identify ineffective teachers, viewing the process as "student success" rather than punitive "evaluation."
System Design and Best Practices (Jason Culbertson)
- TAP Model Components: Combines instructionally focused accountability with multiple career paths (mentor/master teachers) and ongoing professional growth aligned to the rubric.
- Calibration Rigor: Contrasts with other systems where certification takes one day; TAP requires 5-7 days of training, video calibration, and a rigorous online exam (e.g., 74% pass rate in Texas).
- Realistic Distributions: TAP produces realistic bell curves of teacher performance, whereas many state systems (e.g., South Carolina's ADEPT) historically show 98% "satisfactory" ratings due to low rigor.
- Scalability Examples: Tennessee implemented the TAP rubric statewide, achieving record student growth scores after one year despite lacking the full career ladder at that time.
- Resource Efficiency: Utilized online video libraries with short, adaptive modules to support teachers without requiring extensive time away from classrooms.
Implementation Challenges and Consensus Points
- Evaluation Calibration: All panelists agreed that evaluator training must be rigorous to ensure inter-rater reliability and that "compliance-only" evaluations fail to drive improvement.
- Alignment of Metrics: A critical failure point is the misalignment between teacher observation data and student growth data, often caused by state tests measuring rote knowledge while Common Core demands conceptual thinking.
- Labor-Management Relations: Moving from adversarial relationships to collaborative "professional communities" is essential, particularly for resolving the "evaluator vs. evaluatee" barrier.
- Systemic Support: Evaluation systems cannot exist in isolation; they require adequate time, funding (often reallocated from intervention costs), and a supportive school culture.
- Future Outlook: Successful Common Core implementation is dependent on the simultaneous adoption of robust teacher effectiveness systems to support rigorous instruction.