SURE Research · Volume 2  26,544 teachers

Ten years in the classroom. Under two points of improvement.

Across 26,544 teachers in 69 countries, first-year teachers scored 49.5% on classroom competency and teachers past their tenth year scored 51.4% — a correlation of r = 0.03. Experience alone does not repair teaching quality. Coaching does, and coaching has never scaled.

Full reports free to read — no email Methodology & limitations published Figures reusable with attribution
Figure 1 · SURE Research Volume 2

A decade of classroom experience is worth 1.9 percentage points.

80% mastery
First-year teachers
49.5%
10+ years' experience
51.4%
Mean classroom competency score by years of teaching experience
GroupMean score
First-year teachers49.5%
Teachers with 10+ years51.4%
Mastery threshold80%

Base: 20,429 scored attempts with experience data, pooled 2022–2023. Correlation between years taught and competency: r = 0.03. Participation was voluntary, so the population skews motivated.

51%
Mean competency across the whole assessed population.
Pooled 2022–2023 · 46.8% (2022), 53.4% (2023)
0.49σ

Individualised coaching improves teaching practice by 0.49 standard deviations — among the largest effects education research has produced.

60 causal studies · cited in Vol 2
0–5% → 95%

Training alone moves 0–5% of teachers to use a new skill in class. With coaching on their own lessons, about 95%.

Vol 4 · Joyce & Showers
35 pts

The median teacher carries a 35-point spread between their strongest and weakest domain. No single course fits that profile.

Vol 2 · our data
2.4

Formal observations a tenured teacher receives per evaluation cycle — against roughly a thousand lessons taught a year.

Vol 5 · US district average

The gap is specific, not general. Experience does not repair it. Coaching does — and coaching has never scaled.

Read the four findings together and they write a design brief. Development has to begin from an individual diagnosis, because the median teacher's profile is jagged in its own particular way. It has to name one change at a time, because that is what a good coach does. It has to be continuous rather than episodic, because unexamined classroom time does not compound into skill. It has to observe practice rather than self-perception, because the weakest skill cluster in the data is the feedback loop itself.

And it has to break the expert-hours constraint — because coaching's economics collapse when every observation consumes an expert's hour. That last requirement is the one that has been impossible until now, and it is the reason this product exists.

The product, in one move

One recording. Two views. Nothing else asked of anyone.

A teacher records an ordinary lesson on their phone. That recording is the only input the system ever needs. From it, the teacher gets coaching on their own practice, and the school gets a pattern across the staff.

Recording
41:32
Grade 8 · Algebra · ordinary lesson
Class Insight ReportThe teacher's view · private
Strength · clear questioning Signal · wait-time after questions
08:12 Question posed — 1.1s pause before a student is called on
27:48 Worked example paced well; checks for understanding land
Your one thing
Pause 5 seconds after each question to widen student responses.
The school's view · aggregated
Behaviour mgmt
Assessment use
Communication
Aggregate drift across terms — sample
Patterns across staff. Never a ranking, never an individual.

Sample report shown. Scored on the CPAT framework — six domains, 25 sub-domains, the same architecture the research ran on. If something wasn't in the recording, nobody is asked to supply it.

Aggregate patterns for leaders. Individual detail for the teacher only.

How that is enforced →
How it works

From one lesson to a compounding loop.

Step 01

A teacher records

An ordinary lesson, on their own phone. No camera crew, no announced performance — a phone on the desk.

Step 02

SURE listens

The lesson is scored against CPAT's 25 sub-domains — the same instrument the research runs on. Student data is never captured and is always masked.

Step 03

One thing comes back

The teacher's private report names strengths, signals, and a single specific change — the way a good coach ends a debrief.

Step 04

The pattern emerges

Leaders see domain-level patterns across the staff and over time. Never a ranking, never an individual.

Then it repeats — the next recording measures whether the one thing landed. That is the loop observation never closed.

The framework

Scored on CPAT, not on vibes.

Six domains and 25 sub-domains of classroom practice — the same instrument that scored 1.2M+ responses in the research. Every claim in a report names the sub-domain it rests on, so a signal is checkable against the recording, not an opinion to take on trust.

Foundational Explorer Visionary Leader

Every teacher moves along a named progression — one thing at a time.

25 sub-domains · a sample
Learner behaviour management Use of assessment data Effective classroom communication Curriculum planning & progression Differentiated lesson plans Learner needs identification Teaching strategies & methods Reflective thinking skills Purpose of assessments Use of technology to support learning Effective lesson planning + 14 more

Highlighted: the profession's weakest sub-domain in the research — 39.9% correct across 754,971 responses.

The status quo

The observation cycle, side by side.

Annual observation is the instrument schools already pay for. Here is what it delivers, against what one recording habit delivers.

Observation as it existsWith SURE
2.4 formal observations per evaluation cycleEvery lesson a teacher chooses to record
An announced performance, prepared for an audienceOrdinary lessons, as they actually run
One observer's judgement — 0.14–0.37 reliabilityPatterns across many lessons, on one published framework
A verdict at the end of a cycleOne specific next step after each recording
An expert hour consumed per observationNo expert hours — the constraint that kept coaching from scaling

Observation figures from SURE Research Volume 5: US large-district averages; single-observation reliability range from published studies.

SURE Research

Most products cite research.This one is downstream of it.

A standing research programme, not a marketing campaign. We assess classroom competency at scale, publish what we find with the method and the limitations attached, and let anyone check the arithmetic.

26,544
teachers assessed
69
countries
10,562
schools
1.2M+
scored responses
25
sub-domains of practice
How we research

Stated plainly: participation was voluntary, so this population skews motivated and graduate-qualified — which makes a near-50% average the conservative reading, not the alarming one. The method and the limitations ship inside every volume. Figures are free to reuse with attribution. The research mailing list receives research and nothing else.

Browse the research hub → Full reports free to read · PDFs cost an email
Trust

Rules that make recording feel safe.

SURE only works if teachers record willingly. Every rule here exists to protect that.

Teachers own their reports

The individual Class Insight Report reaches the teacher and no one else. Leadership sees aggregates only.

Never a ranking

Leadership views are patterns across staff and over time. No leaderboards, no per-teacher scores, no league tables.

Students are not the subject

Student data is never captured and is always masked. The recording exists to observe teaching, not children.

Nothing beyond the recording

The teacher starts every recording. If something wasn't in it, nobody is asked to supply it — no forms, no self-reports, no uploads.

Privacy law came first

Recording and retention were designed around India's DPDPA, FERPA and GDPR — researched before the feature was built, not retrofitted.

Procurement answers, ready

A procurement pack — DPA, security overview, data-flow map — walked through with your team before any commitment.

Early-access voices

Coaching that finally feels like trust.

My teachers stopped seeing observation as a threat. They see the patterns, I see the patterns, and the detail stays with them. That changed everything.
Priya NairHead of School, K–8
One tap after a lesson and I get my One Thing. No forms, no waiting for an annual review. It's the first coaching that actually fits my week.
Marcus BellSecondary Teacher
Knowing students are never part of the data, and the audio is only ever a recording teachers choose to make, made this an easy yes for our families.
Hana SatoGrade 5 Teacher

Illustrative quotes shown for this early-access preview.

Mapped to the frameworks you already answer to
Ofsted KHDA CIS IB CAEP World Bank Teach
Questions

Asked in every first conversation.

Who sees a teacher's report?
The teacher, and no one else. The architecture says so, not a policy document: the teacher starts every recording, the individual Class Insight Report reaches them alone, and leadership sees domain-level patterns across the whole staff. Never a ranking, never an individual. What teachers own →
What happens to student voices?
Student data is never captured and is always masked. The system observes teaching — questioning, wait-time, feedback, behaviour management — not children.
How much work does this add for a teacher?
One action: record an ordinary lesson on a phone. No forms, no uploads, no self-assessment. The report and the one thing come back without anything else being asked.
Can I trust an AI's read of a lesson?
Treat it as a signal, not a verdict. Scoring runs on CPAT — six domains, 25 sub-domains, published method — and every claim in a report names the sub-domain it rests on, so it can be checked against the recording itself. The research behind the instrument, limitations included, is public.
What does it cost?
Pricing is banded per teacher, per year — no per-recording metering, so nobody is taxed for practising. Pilots are fixed-price and fixed-duration. The full band sheet is shared in the first walkthrough, before any commitment.
What do leaders actually see?
Domain-level patterns across the staff and their movement over time — where the school is strong, where next term's PD should aim. Aggregates only, on the same 25 sub-domains the research uses.
Does it work?
The mechanism is the best-evidenced intervention in the field: individualised coaching moves practice by 0.49σ, and coaching on a teacher's own lessons moves transfer from 0–5% to roughly 95%. Our own product efficacy is a younger claim — we publish numbers with denominators when we have them, and nothing before.
Which frameworks does it align to?
Reports map to the frameworks schools already answer to — Ofsted, KHDA, CIS, IB, CAEP and World Bank Teach — so the evidence lands in language an inspection already speaks.

Decide with the evidence in hand.

The research stands on its own — start there if you like. When you want to see your own classrooms in it, we're twenty minutes away.

Subscribe to the research series.

One email when a volume publishes. No product marketing in it — that is the deal, and it is the reason people stay subscribed across a ten-month decision cycle.

One field. Unsubscribe in one click, any time.

See it on your own classrooms.

Twenty minutes, walked through by someone who can answer methodology questions — not a slide deck.

Or start with a scoped pilot — fixed price, fixed duration.

Pricing is banded per teacher and shared in full in the walkthrough. How pricing works →

SURE

© 2026 SURE · Systematic Upskilling with Results and Evidence. Student data is never captured and is always masked.

PrivacyTerms