Tandem Education
All posts

AI Teacher Evaluation Tools: 8 Questions to Ask Before You Buy

Julian Pscheid

Julian Pscheid

8 min read

Before your district signs a contract for an AI classroom observation tool, get written answers to eight questions. Where does the classroom audio go? What happens to student information in the transcript? Does your data train anyone's model? Who are the subprocessors? Can the tool produce a rating on its own? Can teachers see the evidence? Is AI-written text labeled? And how do you get your data back out? A vendor that can't answer all eight in plain language isn't ready for your teachers.

These questions matter now because districts are buying AI faster than they can evaluate it. In an August Education Week report on AI purchasing, one superintendent summed up the market as “so overwhelming for districts,” and a regional education official put it more bluntly: “Anything with AI is a risky click.” Teacher evaluation is no exception. In March, GovTech reported that administrators are already using AI in evaluations, often on their own initiative and without a district policy, and the AFT's Rob Weil warned that “using it without transparency is a real problem.”

A note on where I'm coming from: I'm the CTO at Tandem, and we build one of these tools. So read this with that in mind. These are the questions I want districts to ask us, and the ones I would ask any vendor if I were sitting on your side of the table.

Why AI observation tools need their own vetting

Most district AI checklists were written with student-facing tools in mind: tutoring bots, writing assistants, adaptive practice. An observation tool is different in two ways. It listens to a room full of children who never signed up for anything. And what it produces feeds into an adult's employment record. Weak answers on privacy put students at risk, and weak answers on judgment put teachers at risk, so the vetting has to cover both.

1. Where does the classroom audio go, and is it ever saved?

This is the first question because it decides most of the others. A tool that records audio and transcribes it later creates a file of children's voices that has to be stored, secured, retained, and eventually deleted. A tool that streams audio to speech-to-text and never writes it to disk creates no such file.

A good answer describes the path end to end: what device captures the audio, where it is sent, whether any service in the chain keeps a copy, and what happens if the connection drops. A red flag is “audio is stored securely,” with no account of why it's stored at all, or a recording you can play back later. Ask your counsel whether your state's consent rules treat stored recordings differently from live transcription, too.

2. What happens to student information in the transcript?

Even if the audio disappears, the transcript stays, and a transcript of a real lesson is full of student names. “Marcus, can you read the next line?” is an ordinary sentence in a classroom and an education record risk in a database. That is why districts should treat any AI observation workflow as FERPA-sensitive.

A good answer is that student names are removed automatically, every time, before anything else touches the text, and that the removal can't be switched off by a busy user. A red flag is a manual redaction step, or a setting the district has to remember to enable.

3. Is our data used to train or improve any AI model?

Ask this about the vendor and about every AI provider the vendor uses. The vendor's own policy may say no while the underlying model provider's consumer terms say something else. Ask which contract governs, and ask to see the clause.

This is quickly becoming a legal requirement as well as a good practice. California's AB 1159, signed in September, bars companies from using student data to train or develop AI models starting next year, and it applies to companies that know their products are used in schools, not only those that serve students primarily. A red flag is any answer that includes the words “anonymized” or “aggregated” without saying exactly what is used, for what, and under which agreement.

4. Who are your subprocessors, and where does our data live?

Every AI tool is a stack of other companies: a cloud host, a speech-to-text service, a language model provider, maybe an analytics vendor. Each one touches your data.

A good answer is a named list, with what each subprocessor does and where it processes data, that the vendor will commit to keeping current in your data privacy agreement. If your state participates in the Student Data Privacy Consortium, start from its standard agreement rather than the vendor's paper. A red flag is “we use industry-leading partners.”

5. Can the tool produce a rating, and who sees it first?

This is the question I'd spend the most time on.

Many AI evaluation tools can take a transcript, map it to a rubric, and propose a score. That sounds efficient. The problem is what happens to the evaluator. When a person sees a confident, well-written AI judgment before forming their own, they anchor on it. We wrote about this automation bias earlier this year: “a human reviews it” is a much weaker safeguard than it sounds when the AI went first. Kim Marshall made a related point in a Fordham Institute commentary, calling the step from AI transcript straight to rubric rating “doubly flawed.”

So don't stop at “does a human approve it?” Ask what the evaluator sees, in what order. A good answer is that the tool presents evidence, the evaluator decides what it means, and any AI suggestion is something the evaluator actively accepts or rejects, item by item. A red flag is a demo that opens on a finished report with ratings filled in.

6. Can the teacher see the evidence used about them?

An evaluation is only as trustworthy as the evidence behind it. One of the strongest arguments for transcription is that it replaces hurried notes and memory with a record of what was actually said, which my co-founder Rolland Hayden argued protects teachers from evaluator bias. That protection only works if teachers can see the record.

A good answer explains what the teacher can see, when, and how they can respond. A red flag is a system in which a teacher's lesson is analyzed by AI and the teacher never sees the analysis. Teachers' unions have good reason to push back on that.

7. Can the evaluator tell which words are theirs?

Six months into using a drafting tool, can an evaluator look at a finished report and tell which sentences they wrote and which the AI suggested? If not, the record stops being a record of professional judgment. It also makes the tool impossible to audit when a teacher disputes an evaluation.

A good answer is that AI suggestions are labeled as suggestions and become the evaluator's text only through an explicit action. A red flag is AI text that blends seamlessly into the draft. In a demo, that seamlessness can look like polish.

8. How do we get our data out, and how do we delete it?

Every contract ends eventually. Ask what format the export comes in, how long data is retained by default, who can trigger deletion, how long deletion takes across every subprocessor, and what proof of deletion you receive.

A good answer is specific and written into the agreement. A red flag is “just contact support.”

How to use these questions

Ask for the answers in writing. Put the eight questions in your RFP or vendor questionnaire, and make the written answers part of the contract. Sales conversations are not commitments.

Ask to be shown, not told. Request a data flow diagram and walk through it with your technology director. Then run a real observation in the pilot and look at exactly what the evaluator sees first.

Bring your teachers' association in before the pilot, not after. AI is already showing up in labor agreements. Detroit teachers' new contract, for example, states that AI cannot replace teacher professional judgment and requires a published list of approved AI tools. Questions 5 through 7 are the ones your teachers will care about most, and answering them together builds the trust that makes the tool useful at all.

Check whether your AI policy covers adults. Many district AI policies were written for student use. If yours says nothing about AI in supervision and evaluation, fix that before you buy.

How we answer them

Elevate streams classroom audio to speech-to-text and never stores it, removes student names from every transcript automatically, runs on enterprise AI terms that prohibit training on customer data, lists its subprocessors publicly, and labels every AI suggestion so the administrator accepts or rejects it. The details are on our security page. For the rest of the list, ask us the same way you'd ask anyone else, and if we can't answer one clearly, count that against us.

Tools that handle the documentation can give principals back hours they should be spending in classrooms. For evaluation, though, the right tool produces evidence and leaves the verdict to the evaluator, and you should be able to see that in how it's built, not only in the brochure.

If you're working through a purchase now and want a second set of eyes on a vendor's answers, including ours, get in touch.

Frequently asked questions

What should a school district ask before buying an AI teacher evaluation tool?

Get written answers to eight questions: where classroom audio goes and whether it is ever stored, what happens to student information in transcripts, whether district data trains any AI model, who the subprocessors are and where data is hosted, whether the tool can produce a rating on its own, whether teachers can see the evidence used about them, whether AI-written text is labeled, and how the district can export and delete its data.

Should AI assign ratings in teacher evaluations?

No. AI can organize evidence and suggest connections to a rubric, but the rating is a professional judgment that belongs to a trained evaluator. When an AI proposes a rating first, evaluators tend to anchor on it, so a tool should keep the evaluator forming their own judgment before seeing any AI suggestion.

Does FERPA apply to AI classroom observation tools?

It can. A transcript of a lesson will almost always contain student names and other student information, so districts should treat AI observation workflows as FERPA-sensitive: use an approved tool under a data privacy agreement, and prefer tools that remove student names from transcripts automatically.

Is it legal to use AI to transcribe a classroom observation?

It depends on your state’s recording and consent laws and on your district’s policies and labor agreements. Some states require consent from everyone whose voice is captured. Ask the vendor whether audio is ever saved as a file, and review the workflow with district counsel and your teachers’ association before a pilot.

Julian Pscheid

Written by

Julian Pscheid

Co-Founder & Chief Technology Officer at Tandem Education

Ready to transform teacher evaluation?