Can You Use AI to Write Teacher Evaluations?
Administrators are already pasting observation notes into ChatGPT, quietly and without a policy. Here is what is legal, what is ethical, and what to say to your staff before someone else says it for you.
By School Solutions | September 2026 | 9-minute read
It's 9:40 on a Tuesday night. Six walkthroughs from last week are sitting in a notebook, three of them in shorthand you're already having trouble reading, and summative conferences start in eleven days. And somewhere in the back of your mind is a question you have not said out loud to anyone in your building: could I just have AI write these?
It's a fair question. The fact that nobody is saying it out loud is the actual problem.
Administrators are already doing this. Not in a district-approved platform with a signed agreement — in a personal ChatGPT tab, at night, pasting in notes that name a teacher and describe their classroom, with no policy, no disclosure, and no idea where that text goes. The silence around it isn't caution. The silence is the risk.
So let's answer the question properly.
The short answer
Yes, you can use AI to write teacher evaluations — with one hard line. You can use AI to organize, structure, and draft evaluation language from evidence you personally collected. You cannot use it to generate the evidence, decide the rating, or supply the professional judgment. The observation is yours. The evidence is yours. The rating is yours. Your signature at the bottom of that document means you made the call, and no tool can make it for you.
Everything else in this article is the detail underneath that sentence.
Is it legal?
The first objection you'll hear is FERPA, and it's almost always aimed at the wrong target.
FERPA governs student education records. A teacher observation write-up is not a student education record — it's a personnel record, governed by your state's employment law, your district's records policy, and your collective bargaining agreement. That distinction matters enormously, because it means the compliance conversation you actually need to have is with your HR director and your union president, not only with your data privacy officer.
Where FERPA does re-enter the picture is at the edges. Observation notes routinely reference students in passing — "two students off task during the opener," "the student with the 504 needed the directions repeated." Those references sit inside a personnel record, which is where they sat long before any AI tool existed. But once you paste that text into a system outside your district, you have made a decision about student-adjacent information, and you should make it deliberately:
Describe student behavior without names. "A student in the back row" carries the same evidentiary weight as a name and creates none of the exposure.
Never paste rosters, IDs, grades, IEP or 504 documents into a general-purpose AI tool. Not as context, not "just to help it understand the class."
Know whether the tool you're using trains on what you type. Consumer chat products and business APIs are governed by different terms, and the default is not the same.
If you want the long version of how this analysis applies to a specific product, our FERPA documentation works through it end to end.
Is it ethical?
Here is the honest version of the ethical question, and it is not really about AI.
An evaluation is a professional judgment about a colleague's practice, attached to their livelihood. What makes it legitimate is not the prose. It's that a trained evaluator was in the room, watched the work, weighed it against a shared standard, and is willing to defend the conclusion in a conference, in a grievance, and if it comes to it, in an arbitration.
Drafting is not judging. When you use a tool to turn "9:05 bellringer on board, 22 of 24 working within 2 min" into a sentence a teacher can actually read and act on, you have not outsourced anything that matters. You've done what you'd do with a template, a colleague's phrasing, or forty minutes on a Sunday — faster.
You cross the line the moment the tool supplies something you didn't observe. A rating you didn't reach. Evidence for an indicator you have no notes on. A "strong classroom culture" you never saw, because the model knows that's what usually goes in that box.
The test is simple: if a teacher challenged a single sentence in this report, could you point to the moment in your notes it came from?
If the answer is yes for every sentence, you're fine. If it's no for even one, that sentence should not be in a personnel file — and that was true before AI, too. The difference is that AI makes it very easy to produce fluent, confident, professional-sounding text about a classroom nobody was standing in. That's the real hazard, and it's a hazard of judgment, not of technology.
What your union will ask — before they ask it
This is the part most administrators skip, and it's the part that ends careers.
In most collective bargaining states, evaluation procedure is negotiated. Introducing a new tool into that procedure without telling anyone is not a technology decision; it's a unilateral change to a bargained process, and it's a grievance waiting for a trigger. The trigger is almost always the same: a teacher receives a report, something in the phrasing feels generic, they ask a question, and they find out from someone other than you.
The fix costs you one paragraph and one faculty meeting. Say it first, say it plainly, and say it before you use it:
"I want you to know how I'm writing observation reports this year. I take the same notes I've always taken, in your classroom, watching your teaching. I now use a tool that helps me organize those notes into the report format — the way I've used templates before. Every piece of evidence in your report comes from something I wrote down while I was in your room. Every rating is my decision. I read every word before it goes to you, and my name is on it. If you ever read something in a report you don't recognize from your own classroom, bring it to me and we'll look at my notes together."
Administrators who say that in September have a non-issue. Administrators who get asked about it in March have a problem. It is genuinely that lopsided.
Two things make that promise real rather than rhetorical: your notes have to actually exist and be retrievable, and you have to mean the last sentence. Offer to show the notes. The offer is almost never taken up, and the willingness to make it is what carries the room.
Where AI helps, and where it must never go
Genuinely useful
Turning shorthand into sentences. The single biggest time cost in evaluation is translation — fragments and timestamps into professional prose. This is mechanical work and it is the work AI is actually good at.
Mapping evidence to the right indicator. If you use Danielson, Marzano, or a state framework, matching what you saw to the correct element is tedious and error-prone at 10pm. A tool that proposes the mapping — and shows you the evidence it used — is a real assist.
Consistency across a caseload. The teacher you observed in October and the one you observed in April deserve the same standard. Nobody's memory delivers that unaided.
Surfacing what you did not capture. The most valuable thing a good tool tells you is that you have insufficient evidence for an indicator — before the conference, while you can still go back in.
Making feedback actionable. "Continue to develop questioning strategies" changes nothing. Turning your notes into a specific, observable next step is a craft skill, and drafting help genuinely improves it.
Off limits
Generating evidence you didn't observe. Fabricated specifics in a personnel file are indefensible in a grievance and dishonest to the teacher.
Producing the rating. A suggested level you review is a starting point. A level you accept without deciding is an abdication.
Filling a form you haven't read. If you sign it, you wrote it. There is no version of this where the tool is at fault.
Anything involving a struggling teacher's improvement plan or a dismissal file without your HR director's explicit sign-off on the process. The stakes are different and so is the scrutiny.
Student data of any kind. Rosters, grades, IEPs, discipline records. A teacher evaluation does not require them, so don't introduce them.
Five questions to ask before you paste anything in
Whether you're evaluating a purpose-built product or deciding whether to keep using a general chatbot, these five answers determine whether you have a tool or a liability.
Does this train on what I type? Consumer chat tiers and business APIs differ. Get the answer in writing, not from a marketing page.
Where is the data stored, and for how long? "In the cloud" is not an answer. A country and a retention period are.
Can it refuse to rate? A tool that always produces a complete report is telling you it will invent evidence to fill a gap. The ability to return insufficient evidence is the single best signal of a product built by someone who has actually sat in a post-observation conference.
Does my district's form survive? If it can't take the form you're required to submit and leave the fields it doesn't have visibly blank, it hasn't solved your problem — it's created a second one.
Who can see my teachers' evaluations? Including the vendor. Ask specifically whether staff can read customer data, and under what circumstances.
Any vendor who gets defensive at question three is worth walking away from. It's the cheapest filter available to you.
Frequently asked questions
Is it legal to use AI to write teacher evaluations?
Yes. Teacher evaluations are personnel records, not student education records, so FERPA is not the governing law. What governs is your state's employment and public records law, your district's policy, and your collective bargaining agreement — which usually means evaluation procedure is a negotiated subject and changes should be disclosed, not assumed.
Do I have to tell teachers I'm using AI?
No law requires it in most states, and you should do it anyway. Disclosure before the fact costs one paragraph in a faculty meeting. Discovery after the fact costs trust you will spend a year rebuilding, and may constitute a unilateral change to a bargained process.
Can AI decide a teacher's rating?
No. A suggested performance level is a draft for the evaluator to accept, change, or reject. The evaluator observed the class, holds the credential, and signs the document — the judgment is theirs and cannot be delegated to software.
Is it safe to put observation notes into ChatGPT?
Notes naming a teacher and describing their classroom belong in a system with a data agreement covering your district, not a personal consumer account. The specific risks are that consumer tiers may retain and train on input, that the account is personal rather than institutional, and that there is no audit trail if the record is later challenged.
Will using AI make my evaluations weaker in a grievance?
Only if you can't trace the evidence. An evaluation is defended by contemporaneous notes and a consistent standard, and a tool that keeps your notes attached to the finished report strengthens that position rather than weakening it. An evaluation containing specifics you cannot source is indefensible — with or without AI.
The part nobody puts in the policy
Most of the anxiety in this conversation is really about something else. It's the worry that if the write-up gets easier, the evaluation becomes less serious — that we're automating the one part of the job that's supposed to be human.
But the write-up was never the human part. Sitting in the back of a room watching a colleague teach, noticing the thing that worked, finding language that makes a teacher want to try something on Monday instead of getting defensive — that's the human part, and it's the part that gets squeezed out first when Sunday disappears into paperwork. We wrote about that tension at length in Stop Grading Teachers. Start Trusting Them.
Getting the four hours back is not the point. What you do with them is.
EduEval is built by a practicing school administrator for the specific problem described here: turning the notes you actually took into a report you can defend. It returns "insufficient evidence" instead of guessing, keeps your district's form intact, and never fills a signature line. If you're comparing options, our overview of teacher evaluation software covers the category honestly, including where we're not the right fit.
About the author
School Solutions
School Solutions writes about instructional leadership, teacher coaching, and AI tools for K–12 schools. Founder of School Solutions, building EduEval — the AI teacher evaluation platform for principals.