Skip to content
All lessons
AI ImpactImpact on Society9-12.AII.IS.1

Three Tests: Bias, Accuracy, Harm

11–1250 minutes90 minutes (extend seminar and writing)

Standard quoted exactly

Evaluate AI-generated output to assess bias, accuracy, and potential harms.

Student-friendly learning targets

  • I can run three separate tests on an AI output: bias, accuracy, and potential harm — and keep the tests from collapsing into one vibe.
  • I can document evidence for each test, including what I still do not know.
  • I can recommend a human action (use, revise, refuse, or escalate) based on the three tests together.

Essential questions

  1. What would have to be true for this output to be biased, inaccurate, or harmful — and how would I know?
  2. Why can a fluent, locally flavored paragraph still fail one of the three tests?
  3. When is the ethical action to refuse an output rather than to edit it?

Objectives

  1. Apply a three-test protocol to at least two teacher-provided AI outputs in different domains (for example, news summary, career advice, health-adjacent general information, or Idaho history).
  2. Distinguish representational bias, statistical bias, and missing-context bias from simple factual error.
  3. Verify or falsify at least two claims with a non-AI source.
  4. Score potential harm by audience and stakes (low inconvenience versus rights, health, reputation, or civic trust).
  5. Write a verdict: use, revise with citation, refuse, or escalate to a human professional.

Key vocabulary

Bias (in output)
A systematic slant in what the model assumes, omits, or emphasizes — for example defaulting to one gender, region, or political frame as normal.
Accuracy
Whether specific claims match checkable evidence from a source that is not the same model.
Hallucination
A fluent claim that is false or unverifiable, including invented citations, dates, or people, presented as if it were known.
Potential harm
Reasonably foreseeable injury if a person acted on the output: medical, legal, financial, reputational, discriminatory, or civic.
Representational harm
Harm from how a group is portrayed or erased, even when no single numeric fact is wrong.
Stakes
What hangs on the decision: a joke caption is low; a scholarship essay, a diagnosis hint, or a news summary used in class is higher.
Refusal
The legitimate academic move of not using an output because it failed a test, rather than polishing it until it sounds fine.

Teacher background

Juniors and seniors have already practiced spotting hallucinations in fluency courses. This lesson is stricter. The standard requires evaluation of output for bias, accuracy, and potential harms as three tests, not as a single dislike. A paragraph can be accurate about Idaho potato acreage and still biased in whose labor it erases. It can be balanced in tone and still invent a statute. It can be true and still be harmful if it offers medical or legal direction a student might follow. Teach a protocol students can reuse in English, government, health, and CTE: (1) who is assumed or omitted, (2) which claims can be checked and against what, (3) who could be hurt if this were trusted. Use printed outputs the teacher generated in advance on a school account, or public examples already in circulation. Do not send student questions about personal health, discipline, or immigration into a live model. Do not use classmates as specimens of bias. Idaho examples help: a summary of a school-board meeting, a wildfire-safety blurb, a Micron-adjacent career paragraph, a history capsule about a treaty or a mining town. Offline fallback is the design, not a backup: dated printouts plus a non-AI reference (textbook page, agency FAQ, newspaper). FERPA: student names do not appear in prompts or in the outputs you file.

Materials and prep

Materials

  • Three-test protocol sheet: Bias / Accuracy / Harm, with evidence and unknown boxes.
  • Packet of four printed AI outputs, labeled A–D, dated, with the prompt shown. Domains: local-news summary, career paragraph, general health lifestyle blurb, Idaho history capsule.
  • Non-AI reference set: one agency FAQ, one textbook or encyclopedia page, one Idaho Education News or local paper excerpt, one official statistics table.
  • Verdict stamps: Use / Revise / Refuse / Escalate.
  • Offline fallback: no live generation during class. If a student wants to test a new prompt, they write it for homework on paper and the teacher decides later whether a school-account run is appropriate.

Before class

  • Generate or collect outputs before class. Include at least one fluent falsehood, one representational slant, and one high-stakes overreach (for example, a lifestyle blurb that sounds clinical).
  • Print references. Do not assume students can search on their phones.
  • Remove any accidental PII from prompts (school names of real minors, staff emails).
  • Block plan: after the protocol, a 20-minute seminar on when refusal is required, then a 20-minute evaluation essay using all three tests.

Instructional sequence

One paragraph, three scores

5 min
  1. Project Output A (a glowing, generic career paragraph about working in Boise tech). Students silently score Bias, Accuracy, and Harm as pass / fail / not enough information.
  2. Do not debate yet. Collect the split: people often pass accuracy because it sounds right.
  3. Tell students the standard names three assessments. Fluency is not one of them.
  4. Reveal one planted error in A after scores are in, to unseat over-trust.

A protocol, not a vibe

10 min
  1. Teach the three tests with definitions and a negative example for each: bias without a false fact; a false fact without obvious bias; a true statement with high harm if followed.
  2. Show how to write evidence: quote the clause, name the missing group or the unchecked claim, name the audience who might act.
  3. Teach verdicts: use (low stakes, checks out), revise (fixable with citation), refuse (do not pass it on), escalate (a professional must own this: nurse, counselor, lawyer, journalist).
  4. Health-adjacent rule: this class escalates; it does not diagnose. Same for legal advice.
  5. FERPA: we evaluate teacher-provided text. We do not paste a classmate’s essay into a model to grade its bias.

Work Output B together

10 min
  1. Read a printed wildfire or air-quality blurb that mixes good public-safety language with an invented evacuation zone or a wrong agency name.
  2. Fill the protocol as a class. Accuracy fails on the invented zone. Bias may appear if only valley cities are mentioned and reservation or rural communities are omitted. Harm is high if someone drove toward a fake zone.
  3. Model checking the non-AI reference (agency FAQ). Date the check.
  4. Choose a verdict. Prefer refuse or revise over use. Require a reason that cites a test, not I wouldn’t trust AI.

Two outputs, two verdicts

12 min
  1. Pairs take Outputs C and D. Complete a full three-test sheet for each and issue a verdict.
  2. At least one claim per output must be checked against the paper reference set. If the reference is silent, they must write unknown — cannot verify here.
  3. A pair may not issue the same verdict for both outputs without a specific argument; the packet is designed to differ.
  4. Circulate to stop collapsed scoring (all three tests marked fail because the student dislikes chatbots).

Real-world examples

  • A chatbot summary of a school-board meeting that invents a vote tally.
  • Career advice that assumes a four-year degree is the only path into Micron-adjacent manufacturing or a clinic.
  • A history capsule that describes a treaty solely from a settler newspaper’s voice.
  • A wellness paragraph that edges into dosage or diagnosis language a student might follow.
  • Image-model captions that misidentify a tribal event as a costume party (link to representational harm).

Hands-on activity

Verdict wall

8 min
  1. Pairs post only their verdict and the one test that drove it.
  2. If two pairs disagree on the same output, they have 90 seconds to cite a line, not a feeling.
  3. Teacher highlights any pair that used unknown honestly.
  4. Collect protocol sheets.

Discussion questions

  1. Can an output fail the harm test even if it passes accuracy? Give a line from today’s packet.
  2. Who is the audience that makes a wildfire error more serious than a sports-recap error?
  3. When should a student escalate instead of revising?
  4. Is representational harm a bias test, a harm test, or both? Why split them on the sheet?
  5. What non-AI source would you trust in this building, and what are its limits?

Differentiation

Support

  • Provide a completed protocol for Output A as a model. Students complete C with sentence frames.
  • Highlight the checkable claims in C and D so verification is findable.

Challenge

  • Write a revised version of a refused output that would pass all three tests, with citations to the paper references.
  • Design a fourth test (for example, provenance) and argue whether the standard already covers it.

Multilingual learners

  • Allow annotation of bias in how the output treats additional-language speakers.
  • Provide the protocol headings with student-built glosses; the evaluation remains in academic English with support.

IEP / 504

  • Offer a larger-print protocol and the option to complete one output in depth.
  • Oral verification with the teacher using the paper reference is acceptable for the accuracy test.

Assessment

Formative

  • Warm-up split scores and guided-practice protocol.
  • Unknown boxes used rather than guessed.

Summative

  • Two completed three-test sheets with verdicts, scored on separation of tests, evidence, and an appropriate escalate/refuse when stakes are high.
  • Block extension: 300–400 word evaluation of one output using all three tests and a named non-AI source.

Success criteria

  • Treats bias, accuracy, and harm as separate judgments with evidence.
  • Checks at least one claim against a non-AI source or marks it unverified.
  • Issues a verdict that matches the stakes, including refuse or escalate when warranted.

Responsible use, ethics, and privacy

Responsible use

Teacher-provided outputs only during the period. No personal medical, legal, or disciplinary questions submitted to a model. No classmate work used as a specimen. Offline print is the lesson. If a district tool is used later, the teacher runs it on a school account.

Ethics

Evaluation is a civic skill: fluent text can still smear a group, invent a fact, or put a person in harm’s way. Refusal is an ethical action, not a lack of tech-savvy. Students should not be rewarded for being merely suspicious of all AI, nor for being merely impressed.

Privacy

Do not put student names, health questions, or discipline stories into prompts. FERPA covers education records; health and counseling topics belong with professionals, not with a chatbot. Filed student evaluations of the packet should not include personal data.

Reflection

  1. Which test was hardest to keep separate from the others?
  2. When did fluency almost talk you out of a fail?
  3. What source in this building will you use the next time a claim looks convenient?

Homework

On paper, apply the three tests to a printed output your teacher sends home (or to a screenshot you already have — do not generate a new one). Check one claim with a library or official site, not with another chatbot. Bring the marked sheet.

Closing

Read one refuse and one escalate verdict aloud. The standard is evaluation, not mood. Next class asks how workers in Idaho already use these tools to solve problems — and how they still apply tests like yours.

Extensions and cross-curricular links

Go further

  • 90-minute block: seminar on refusal, then a formal evaluation essay.
  • Apply the protocol to a student-chosen output from a district-approved tool, with teacher-run generation only.
  • Compare two models on the same prompt (teacher-run, printed) and evaluate whether disagreement is evidence.
  • Bring the protocol into a government or English paper as a methods paragraph.
English Language Arts
Rhetorical analysis of fluent emptiness; citation versus invented sources.
Government / journalism
Meeting summaries, vote tallies, and civic harm when the record is wrong.
Health
Why lifestyle blurbs that sound clinical are escalated, not edited, in class.
History
Whose voice is default in a capsule about treaties, labor, or towns.