
How Geena Davis's AI Tools Measure Media Representation
This article explains how the Geena Davis Institute uses GD-IQ and Spellcheck for Bias to quantify screen time, speaking time, and stereotypes in film and TV, and how major studios like Disney and NBCUniversal have adopted these tools to inform their content.
Updated:
Representation in film and television is easy to argue about and hard to settle if the evidence stays at the level of impressions. One viewer says women seem underwritten. Another says a cast looks diverse enough. A studio can point to a few prominent characters. What changes the conversation is a tool that can ask smaller, colder questions: who is on screen, for how long, who speaks, and whose traits are attached to what kinds of roles.
That is where Geena Davis's activism became especially consequential for media impact research. The Geena Davis Institute did not only campaign for better representation; it helped build instruments for measuring it. Its two most important AI-related tools work at different moments in the production cycle: GD-IQ studies completed audio-visual media, while Spellcheck for Bias examines scripts before they become filmed scenes.

Two Tools, Two Points In The Workflow
The distinction matters. A finished film can reveal patterns that no one noticed while watching casually, but by then the casting, editing, and dialogue decisions are already embedded in the work. A script tool operates earlier, when characters can still be rewritten, scenes can be rebalanced, and stereotyped descriptions can be questioned before production costs make change harder.
| Tool | Material Analyzed | What It Measures | Why It Matters |
|---|---|---|---|
| GD-IQ | Finished film and television content | Screen time and speaking time using facial detection and voice recognition | Turns visible and audible presence into measurable evidence |
| Spellcheck for Bias | Scripts before production | Representation gaps and stereotypes across gender, race/ethnicity, disability, and LGBTQIA+ status | Lets creators examine bias while the work is still changeable |
This is a useful division for students because it separates two research problems that often get blended together. One problem is descriptive: what appeared in the final media text? The other is diagnostic and intervention-oriented: what signals are already present in the script that may produce unequal representation later?
GD-IQ Measures What Viewers Usually Estimate
GD-IQ, short for Geena Davis Inclusion Quotient, was developed with the University of Southern California's Signal Analysis and Interpretation Laboratory after a $1.2 million Google grant in 2012. The tool uses automated audio-visual analysis, including facial detection and voice recognition, to measure character screen time and speaking time down to the millisecond.[1]

The millisecond detail is not a decorative technical claim. Manual coding can produce valuable research, but it is slow, labor-intensive, and usually limited by sample size. An automated system can process much larger collections and apply the same measurement rules across them. That does not make every judgment automatic or neutral, but it changes the scale at which a cultural pattern can be observed.
GD-IQ's early findings showed why this mattered. In films from 2014 and 2015, male characters appeared on screen approximately twice as often as female characters. The imbalance was not only about who was cast; it was about how much of the film's time those characters occupied.[2]
The lead-character comparison is especially instructive. When the lead character was male, he appeared on screen and spoke about three times as much as when the star was female.[2] That finding sharpens a common but vague complaint: a film can have a female star and still allocate its visual and verbal attention very differently from the way it treats a male-led story.
From Counting Presence To Seeing Scale
A later Google partnership expanded the scale further through MUSE, an AI system used to study 12 years of television programming. The analysis covered 440 hours of TV and processed more than 12 million second-by-second face appearances in under 24 hours.[3]
Those numbers are not just impressive throughput. They indicate a different kind of media study. Instead of asking a small team to code a narrow sample, researchers could examine second-by-second appearance patterns across a much larger time span. The unit of analysis becomes finer, while the collection becomes broader.
The findings did not point to a simple story of solved gender representation. In 2021 programming, male characters still received about 16% more screen time than female characters. Women over 60 received less than 1% of screen time.[3] That second figure is the kind of result that tends to disappear in broad diversity language. Gender parity in the aggregate can still leave older women nearly absent.
The Institute also reports persistent gaps for disabled and LGBTQIA+ characters, with disability representation at 2.5% and LGBTQIA+ representation at 1.1% in the cited context.[2] These are narrow measurements, not full explanations of social experience. But they make absence harder to wave away as a matter of personal perception.
Spellcheck For Bias Moves The Question Earlier
GD-IQ reads completed media. Spellcheck for Bias reads scripts. That shift changes the practical stakes of the measurement. A script-level tool can flag patterns before actors are cast, scenes are shot, and edits lock in the distribution of attention.
Spellcheck for Bias analyzes scripts for stereotypes and representation gaps across gender, race and ethnicity, disability, and LGBTQIA+ status.[4] Its name is plain, but the analogy is useful: just as a writing tool can flag a spelling problem before publication, this tool is meant to flag representational patterns before they become screen culture.
The important methodological point is that a script tool cannot measure final screen time. It can evaluate character descriptions, dialogue distribution, identity categories, and stereotype signals in the written material. That makes it less like a census of the finished product and more like an early warning system inside development.
Disney began using Spellcheck for Bias in 2019.[4] Universal used it in 2020 for Latinx representation, and NBCUniversal expanded its use in 2021 for Latino, Black, and AAPI representation analysis.[5] Those cases matter because they place the tool inside major content organizations rather than only in academic demonstration settings.
Adoption Is Evidence Of Use, Not Proof Of Repair
Studio adoption should be read carefully. It shows that large entertainment companies found the tools usable enough to bring into content workflows. It does not, by itself, prove that the resulting films and shows became equitable, or that every flagged issue was changed.
The Institute's 2024 impact study offers a useful but limited signal. In a survey of its own industry partners, 94.5% said the Institute is impactful, 79.2% said they had used GDI research to inform their work, and 97.5% said they would recommend GDI resources. Nearly half reported integrating GDI insights into more than 50% of their projects.[5]
Those figures are best understood as evidence of partner uptake and perceived usefulness. Because they are self-reported by the Institute's partners, they should not be treated as independent outcome measurement. A stronger impact claim would need to connect tool use to documented changes in scripts, casting, editing, release patterns, or audience-facing representation over time.
What The Tools Teach About AI And Cultural Research
GD-IQ and Spellcheck for Bias are useful examples because they do not ask students to treat AI as magic. Each tool is tied to a research design question. What counts as presence? Is screen time enough, or does speaking time matter too? Which identity categories are included? Are older women visible in the data, or hidden inside a broad gender total? Is the tool examining a finished artifact or an editable script?
The Institute has also made an open-source skin tone classifier available based on the Monk Skin Tone Scale.[2] That detail belongs in the same methodological conversation: representation research depends not only on having more data, but on deciding which categories and visual distinctions a tool is built to recognize.
The strongest claim is also the most careful one. These tools matter because they turn media representation into measurable evidence at a scale manual coding could not easily reach, and because major studios have used them in production-related workflows. They do not show that representation problems are solved. They show that some of the patterns are now countable, comparable, and harder to dismiss.
For students studying AI, media, or social impact, that is the lesson worth keeping. Machine learning can widen the field of view, but the quality of the judgment still depends on what is measured, whose categories are included, and what institutions do after the results arrive.
References
- Geena Davis Institute on Gender in Media, Wikipedia.
- About Us, Geena Davis Institute.
- Using AI to study 12 years of representation in TV, Google Blog.
- Disney is using AI developed by Geena Davis to correct gender bias, Fast Company, 2019.
- Impact Study, Geena Davis Institute, 2024.
Comments
Join the discussion with an anonymous comment.