The Dataset Decides
Standard quoted exactly
Analyze how choice of data sets used to train AI models can lead to potential bias in the output.
Example from the standards. A camera's face recognition tool only works on people with specific skin pigments due to input/training data.
Student-friendly learning targets
- I can explain how the mix of examples in a training set can cause biased output.
- I can use the face-recognition example to show what happens when some skin tones are missing.
- I can redesign a small Idaho-relevant dataset so it covers people, places, or cases that were left out.
Essential questions
- If the output is biased, what was missing from the data?
- Who gets to be in the training set, and who never appears?
- How would we build a fairer set for an Idaho job without collecting private student data?
Objectives
- Students will trace a biased output back to a specific gap in a training set, using the standards face-recognition example.
- Students will inventory a printed mini-set and list who or what is overrepresented, underrepresented, or absent.
- Students will propose additions and non-examples that would make an Idaho model less brittle.
- Students will distinguish 'more data' from 'better, more representative data.'
- Students will state a privacy rule: we do not fix a dataset by scraping classmates' faces or records.
Key vocabulary
- Dataset
- A collected set of examples used to train or test a model, such as photos, sensor logs, or labeled sentences.
- Training data
- The portion of examples the model learns from. Gaps and slants here often reappear as biased output.
- Representativeness
- Whether the examples cover the people, places, and conditions the model will actually face.
- Underrepresentation
- When a group or condition appears too rarely in the data for the model to handle it well.
- Sampling
- The choices about which examples are collected, kept, or thrown out before training.
- Generalization
- How well a model works on new cases. It fails when new cases look unlike the training set.
- Face recognition
- A system that matches or identifies faces from images. Accuracy can collapse for skin tones, lighting, or ages that were scarce in training photos.
- Non-example
- A labeled case of what something is not, which helps a model learn boundaries instead of a single default.
Teacher background
ECT.3 asked students to spot biased output. ECT.4 asks them to look one step upstream: the dataset. The standards example is blunt. A camera's face recognition tool only works on people with specific skin pigments because of input and training data. Ninth- and tenth-graders can grasp that without building a model. If the photos are mostly light skin in office lighting, darker skin, harsh sun on a fireline, or a welder's hood shadow will fail. The same story plays in Idaho systems that are not about faces. A wildfire model trained on California chaparral will misread Payette timber. A yield model trained on Iowa corn will stumble on Magic Valley potatoes and sugar beets. A hospital sepsis alert tuned on coastal research hospitals will misfit a 25-bed critical-access site. Teach one rule of thumb: the dataset decides who the system can see. Teach a second rule: students must not 'fix' a set by photographing classmates or uploading FERPA-protected records. Use printed cards, a projected demo of a failed match, and a paper redesign. Keep the math light. The civic point is heavy enough.
Materials and prep
Materials
- Projector and a teacher-made slide of two training envelopes: Set A (narrow) and Set B (broader)
- Printed fake-output packet: face-recognition example, Idaho mini-sets (wildfire, potatoes, hospital), and blank redesign sheets
- Card decks of labeled examples and non-examples (clip art or public photos, no student photos)
- Sticky notes for 'who is missing'
- Optional Chromebooks; paper is sufficient
- Privacy reminder poster: no classmate faces, no real medical records
Before class
- Print cards that never include student or staff photos. Use public-domain or drawn faces spanning skin tone, age, and lighting.
- Build three Idaho mini-sets with obvious gaps: California-only fires, Iowa-only crops, big-city-only hospitals.
- Save a screenshot of a failed match or a mismatched wildfire map as the offline demo.
- Check district policy on any live face-tool demo; default to paper if unsure. Do not run recognition on people in the room.
Instructional sequence
Two envelopes
5 min- Hold up Envelope A (ten similar light-skin office photos) and Envelope B (mixed tone, age, outdoor light). Ask which 'camera tool' will fail on a firefighter in ash, and why.
- Students write a one-sentence prediction on the packet.
- Reveal the standards example. Connect envelope A to 'only works on people with specific skin pigments.'
Gaps in, failures out
10 min- Diagram: sampling choices to training set to model to output. Put a hole in the set and a matching failure in the output.
- Teach the face-recognition example in the standards wording, then translate: missing tones and lighting become missed or mislabeled people.
- Map three Idaho twins on the board: California fire data vs. Idaho timber; Iowa corn vs. potatoes; coastal hospital vs. rural hospital.
- State the privacy line: we study this with public cards, never with the class as a dataset.
Inventory a mini-set
8 min- As a class, inventory the wildfire mini-set: fuel type, slope, state, season. Tally what dominates.
- Students place sticky notes: Overrepresented / Underrepresented / Absent.
- Predict one biased output (for example, underestimating timber crowning, or ignoring irrigation canals as fire breaks).
Choose an Idaho set to repair
10 min- Each student picks potatoes, hospitals, ski-area demand, or trades-safety photos.
- On the redesign sheet they list five current examples, three missing examples, and two non-examples.
- They write four sentences: how today's set would bias output, what they would add, and what they will not collect (student faces, patient charts, home addresses).
Real-world examples
- Face tools at a concert or school event can fail on darker skin if the training photos were narrow, which is the standards example in public.
- Wildfire models trained on Southern California brush misfit Payette and Boise National Forest timber and Idaho's canal landscape.
- Ag models trained on Midwest corn miss potato, sugar beet, and trout-farm signals in the Magic Valley.
- Hospital early-warning scores tuned on large coastal hospitals can alarm wrong in a rural Idaho critical-access hospital.
- A recreation demand model trained on Colorado ski data will misread Bogus Basin weekday patterns and local school calendars.
Hands-on activity
Rebuild the deck
12 min- Pairs receive a biased deck (faces, fires, or crops) and must add public cards to improve representativeness. They also add two non-examples.
- They write a 'dataset label' as if they were handing the set to a county IT shop or a Micron intern: source, who is included, known gaps, privacy notes.
- Teacher may project a saved failed-match demo. No live scanning of anyone in the room. If Chromebooks are on, students only open the teacher packet PDF.
- Two pairs trade decks and try to name a remaining gap in 60 seconds.
- Paper packets are the full activity if cards cannot be cut; students circle additions on a printed menu of public examples.
Discussion questions
- Why is 'just add more data' not always a fix?
- Who should have a say before a school or county trains a face or license-plate tool?
- How could a potato model be biased even if every photo is technically 'accurate'?
- What is the difference between a missing group and a missing condition (night, smoke, winter)?
- Why is using the class as a face dataset a FERPA and ethics failure even if the lesson is about fairness?
Differentiation
Support
- Give a filled inventory table with one column already completed.
- Offer a menu of missing examples to circle rather than invent.
- Allow a poster diagram instead of the four-sentence write.
Challenge
- Argue when a local-only set can also be biased (too small, too one-farm, too one-season).
- Add a non-example that prevents a dangerous confusion (smoke vs. fog, bruise vs. dirt).
- Connect to hiring or lending tools in a short risk paragraph, still without collecting PII.
Multilingual learners
- Clarify pigment, representativeness, and sampling with visuals before the inventory.
- Note that language datasets can exclude home languages the same way photo sets exclude skin tones.
- Allow dataset labels to be bilingual.
IEP / 504
- Pre-cut cards and limit the rebuild to adding three stickers from a menu.
- Provide a scribe or speech-to-text on a school account that does not leave the district.
- Reduce independent practice to one mini-set with sentence starters.
Assessment
Formative
- Warm-up prediction connecting Envelope A to the standards example.
- Sticky-note inventory during guided practice.
- Redesign sheets listing missing examples and a privacy refusal.
Summative
- Exit ticket: Quote the face-recognition example, then explain in 5–8 sentences how a dataset choice caused the failure and how an Idaho twin (fire, farm, or hospital) could fail the same way.
- Dataset label scored for source, gap, proposed fix, and privacy constraint.
Success criteria
- I connected biased output to a specific gap in the training set.
- I used the face-recognition example accurately.
- I proposed better examples and non-examples, not just 'more data.'
- I refused to use student PII or classmate photos as a fix.
Responsible use, ethics, and privacy
Responsible use
Study datasets with public, teacher-provided cards. Do not scrape the web for faces. Do not photograph the class. Do not download a real recognition app onto student phones for this lesson.
Ethics
People who were missing from the data are the people most likely to be misidentified, denied, or ignored. Building a fairer set is a justice task, not a cosmetics task. Consent and dignity beat a slightly better accuracy score.
Privacy
FERPA and basic ethics: no student PII, no classmate or staff face scans, no real patient records, no home addresses in a 'local dataset.' If a district tool is used, stay on the teacher prompt and public samples. Printed packets are the default.
Reflection
- Which missing group or condition would have hurt someone I know?
- What will I ask the next time a vendor says their model 'works for everyone'?
- How will I remember that more data is not the same as representative data?
Homework
On paper, invent a tiny training set (eight labeled examples) for one Idaho task: identifying ripe vs. unripe fruit, spotting a dull weld, or flagging a full campground. List two missing examples and one thing you will not collect because it is private. No photos of people you know.
Closing
Hold Envelope A and Envelope B again. Students finish the stem: 'The dataset decides ___.' Collect redesign sheets. Next hallucination lesson will show a different failure: the model can also invent what was never in the set.
Extensions and cross-curricular links
Go further
- 90-minute block: students write a one-page data card for a fictional county tool (park cameras, irrigation alerts) including who must consent.
- Statistics: sample vs. population using Idaho counties, not faces.
- Biology / environmental science: why California fuel models transfer poorly to Idaho timber.
- AITA preview: later technical courses will clean and split datasets; this lesson stays on civic judgment.
- Statistics
- Sampling bias and whether a sample represents the population the model will face.
- Civics
- Public cameras and recognition tools require policy, not just accuracy claims.
- Agriculture
- Local crops, irrigation, and pests must appear in training if a model will run in the Magic Valley.
- Health Sciences
- Clinical scores fail when training hospitals look nothing like rural Idaho sites.