Critique the Corpus
Standard quoted exactly
Evaluate training data by examining its source, quality, representativeness, potential biases, and privacy implications.
Example from the standards. Pick a dataset commonly used to train a model like ImageNet and critique it on the listed criteria.
Student-friendly learning targets
- I can critique a well-known training corpus (ImageNet) on source, quality, representativeness, bias, and privacy.
- I can apply the same five criteria to an Idaho public dataset (yield, gauges, or wildlife photos).
- I can recommend a go / use-with-cautions / do-not-use verdict with evidence.
Essential questions
- Who took these photos or measurements, who is missing, and who never consented?
- If ImageNet trained a generation of vision models, what did those models inherit?
- When is an Idaho public dataset still a bad training set?
Objectives
- Close-read a teacher-provided ImageNet fact sheet (source, size, labeling method, known issues) — not a live download of the corpus.
- Score ImageNet on the five criteria with evidence, not vibes.
- Score one Idaho public set on the same rubric.
- Write a verdict a district tech director or extension agent could use.
Key vocabulary
- Corpus
- A collected body of training items (images, text, tables). ImageNet is a famous image corpus.
- Source
- Where the items came from and who funded, scraped, or labeled them. Source is the first critique, not the last.
- Representativeness
- Whether the corpus looks like the world the model will face (geographies, people, seasons, devices), not just the world that was easy to scrape.
- Label quality
- Whether the tags are accurate, consistent, and at the right grain (a 'apple' label on a phone logo is a quality miss).
- Privacy implication
- Whether people in the data could be identified, were asked, or can get out; also whether locations reveal homes or tribal resources.
Teacher background
Students critique ImageNet as the standards example, using a fact sheet you provide (origin at Stanford/Princeton, millions of labeled web images, WordNet synsets, known problems: western/English web scrape, offensive synsets later removed, people photographed without meaningful consent, label noise, underrepresentation of many geographies). Do not download ImageNet in class; it is huge and messy. Pair it with a local public analog: iNaturalist research-grade observations in Idaho (source and consent differ from a web scrape), USDA NASS tables (survey, not photos), or USGS gauges (instruments, not faces). The five criteria are the rubric: source, quality, representativeness, potential biases, privacy. DA.5 was 'fix the mix'; DA.7 is 'should this corpus train anything at all?' Spreadsheet scoring 1–4 per criterion is enough.
Materials and prep
Materials
- ImageNet one-pager: origin, how labels were made, approximate scale, 3 documented critiques (people without consent, geographic skew, label noise/offensive categories). Citations to public writeups, not leaked files.
- Idaho analog one-pager: pick iNaturalist Idaho observations or USDA NASS potato by county or USGS Snake River daily.
- Five-criteria rubric (source, quality, representativeness, potential biases, privacy) with 1–4 anchors.
- Verdict sheet: go / use with cautions / do not use for X task.
- Optional printed sample label list (public synset names), no student photos.
Before class
- Confirm analog dataset is public and has no student or easily identified private farm photos. iNaturalist: research-grade, obscuring of threatened species already in the source.
- Do not ask students to search for random ImageNet images of people.
- Pre-score both sets on the teacher key so you can coach evidence ('your representativeness 4 needs a citation from the one-pager').
Instructional sequence
Would you train on this shoebox?
5 min- Show a shoebox metaphor: 14 million photos grabbed from the web and tagged by workers. Fast first impressions: source? consent?
- Reveal the five criteria on the board. Today every claim needs a criterion name.
Five questions every corpus must survive
10 min- Source: who collected, scraped, paid, labeled? Quality: noise, duplicates, wrong tags. Representativeness: geography, time, devices, people. Bias: systematic skews (skin tone, language, crop variety, season). Privacy: faces, homes, GPS, kids, tribal and sensitive locations.
- Walk ImageNet through source and privacy first using the one-pager. Do not sensationalize; be specific.
- Note: a corpus can be famous and still fail a criterion. Fame is not a fifth star.
Score ImageNet as a class
12 min- Read the one-pager silently for 3 minutes. Pairs draft a 1–4 score for source and for privacy with a quoted phrase as evidence.
- Share two scores. Disagreement is useful if evidence is cited.
- Complete quality, representativeness, and bias as a class so the method is visible.
Idaho analog on the same rubric
10 min- Pairs score the analog dataset on all five criteria with evidence from its one-pager.
- They write a verdict for a stated task (e.g., 'train a statewide bird classifier' vs. 'publish a county yield nowcast').
- Offline: rubric on paper; no need to open the live dataset portal.
Real-world examples
- ImageNet trained a decade of vision models; inherited web-scrape geography and consent failures show up in production cameras.
- iNaturalist Idaho: better consent norms and research-grade filters, but road-accessible locations and charismatic species still skew representativeness.
- USDA NASS: strong source documentation, but small counties may be suppressed for privacy — a quality/representativeness tension.
- USGS gauges: excellent source and quality, weak representativeness if you train a 'all Idaho streams' ice model on the Snake alone.
Hands-on activity
Two-corpus briefing
8 min- Each pair writes a 6-line briefing: ImageNet verdict for 'general object recognition in Idaho schools'; analog verdict for its stated task.
- They must name the weakest criterion for each.
- Two pairs swap and try to overturn one score with evidence from the one-pager only.
Discussion questions
- Should a school ever fine-tune a vision model on ImageNet-derived weights without reading this critique?
- NASS suppresses small-county data to protect farms. Is that a privacy win, a representativeness loss, or both?
- If a corpus is 'the best we have,' does that make it acceptable?
- What would you demand from a vendor who says 'we trained on a large public dataset' and will not name it?
Differentiation
Support
- Rubric with sentence starters and highlighted evidence lines on the one-pager.
- Score only three criteria (source, representativeness, privacy) if five is overload.
Challenge
- Compare ImageNet to a second famous corpus fact sheet (e.g., a large language-text crawl described, not downloaded) on privacy.
- Propose a replacement sampling plan for an Idaho crop-disease corpus that would pass all five at 3+.
Multilingual learners
- Criteria names with plain glosses; students may write evidence quotes in English (from the sheet) and commentary in another language.
- Discuss how English-only labels in a corpus are a representativeness issue.
IEP / 504
- One-pager in large print; oral verdict with the five names on a checklist.
- Partner reads; student owns two criteria.
Assessment
Formative
- Paired ImageNet scores for source and privacy with a quoted phrase.
- Teacher listens for criterion names, not just 'it's biased.'
Summative
- Completed two-corpus rubric plus verdicts. Each score has evidence from the fact sheet.
- A rant without the five criteria does not meet the standard.
Success criteria
- Student addresses all five listed criteria for at least one corpus.
- Student uses source evidence, not only reputation.
- Student names a privacy implication that is not 'hackers' — consent, identification, or location.
Responsible use, ethics, and privacy
Responsible use
Critique from fact sheets. Do not download ImageNet, do not scrape faces, do not open random synsets of people in class.
Ethics
Famous datasets normalized taking without asking. Students should leave able to say no to a corpus, not only how to fine-tune it.
Privacy
FERPA plus ordinary privacy: no student photos as 'our ImageNet.' For Idaho analogs, prefer gauges and county-level surveys over identifiable homesteads. If a photo set includes people, stop and switch packets.
Reflection
- Which of the five criteria was hardest to score, and what evidence did you wish you had?
- How would you explain ImageNet's privacy problem to a principal in two sentences?
- What corpus will you refuse to treat as 'neutral' after today?
Homework
Using only the take-home one-pagers, write a one-page critique of ImageNet on the five criteria and a half-page critique of the Idaho analog. End each with a verdict for a named task. No image downloads.
Closing
Source, quality, representativeness, bias, privacy — five doors a training set must pass. ImageNet is a caution, not a mascot. Public Idaho tables can fail too. That is the last Data and Analysis move: decide whether the corpus deserves a model at all.
Extensions and cross-curricular links
Go further
- 90-minute block: a third corpus (energy load public tables) and a gallery of verdict posters; or a letter to a vendor asking the five questions.
- Python extension: summarize column missingness and county counts on the NASS CSV as quality/representativeness evidence — still no student PII.
- Card sort of headline claims ('1 million images!') onto the five criteria they actually address (often none).
- Tie back to DA.5: augmentation cannot fix a corpus you ethically should not have.
- Civics / law
- Consent, public-records, and suppression rules (NASS) are policy, not only data science.
- Art / media
- Web-scraped photos have photographers and subjects; a corpus is a pile of other people's work.
- Biology
- iNaturalist quality grades and threatened-species location hiding are field-biology ethics.