We took a close look at Alltold, the company measuring how people are represented across ads, TV, film, and now generative AI output, with identity scales co-created with advocacy groups like GLAAD and RespectAbility. Inside: why fifty years of Super Bowl ads went unmeasured, the wedge of building measurement methodology with the communities being measured, which enterprise buyers this actually lands with, and why most "content intelligence" tools are optimizing for the wrong number.
Somewhere between 1967 and 2025, a half-century of Super Bowl ads happened, and almost nobody could tell you who was actually in them. Not anecdotally. Not with numbers. Studios spend nine figures casting people with extraordinary care, then hand that same job to a generative model trained on scraped data, and wonder why their teams quietly pledge to avoid GenAI for any content with people in it. That is the gap Alltold set up shop in.
The buyers Alltold serves are not short of conviction. Brand and marketing analytics teams, inclusion teams at consumer brands, and studio research groups all believe, with reasonable evidence, that how people show up on screen moves the numbers: sales, viewership, ratings. What they lack is the measurement. Representation has lived in the land of vibes and annual reports assembled by interns with spreadsheets, which means it never had a seat at the budget table. GenAI platform teams have a sharper version of the same ache. Their enterprise customers want guardrails, bias reporting for datasets and models, and some proof that generated people will not embarrass the brand. Without instrumentation, "responsible AI" stays a slide, not a system.
The wedge: measurement built with the people being measured
Plenty of companies slap a demographic classifier on video frames and call it representation analytics. Alltold's differentiation is methodological, and it is genuinely hard to copy. Their identity attribute scales were co-created with advocacy groups and academics, including GLAAD and RespectAbility, and the design choices show it. They annotate gender expression rather than binary gender. They observe the sexual orientation of interactions rather than labeling individuals. This is not political nicety; it is measurement honesty. As they put it themselves: are we seeing two female friends playing with one of their babies, or a lesbian couple playing with their baby? That annotation is hard for humans and models alike, but as they argue, getting some fraction wrong because the task is hard is categorically different from getting it wrong because your model leaned on stereotypes.
The pedigree backs the method. CEO Morgan Gregory came through Google Cloud OCTO, Google Research, BCG, and an MIT MBA; CTO Kree Cole-McLaughlin spent eight-plus years at Google building Responsible AI media models. This is a team that has already done the unglamorous version of this work inside a company famous for refusing to ship the glamorous version too early. And critically, they publish. The quarterly benchmarks with Innovid measure the most-served ads each quarter, top 500 by impressions, with thousands of detected people per report (Q3 2024 alone: 2,894 people detected). Their Super Bowl retrospective covers fifty years of ads across a multi-year report series. Public benchmarks give buyers a comparison set, which almost no vendor in this space is willing to provide, because most incumbents sell private dashboards that can never be checked against anything.
The ICP they actually win: enterprises with a mandate and something to lose
This is not a self-serve tool for mid-market marketers. Three segments land hardest. First, consumer brands and their creative agencies with formal inclusion and brand-safety mandates, where marketing analytics teams need campaign-level representation data that survives a CFO's questions. Second, TV and film studios analyzing representation across shows, seasons, and whole libraries, a corpus-scale problem no human team can tally. Third, and the fastest-moving: generative AI image and video platforms that need dataset and model bias evaluation plus enterprise guardrail tooling, because their buyers are the very brands in segment one. The common thread is enterprise scale plus reputational exposure. A platform selling image generation to a Fortune 500 brand cannot answer "why do all your CEOs look the same" with a shrug. Alltold sells the instrumentation that turns that question into a report.
What the category still gets wrong: optimizing for the label, not the person
The broader people-analytics and content-intelligence category optimizes for classification accuracy on coarse labels: man/woman, old/young, skin-tone bucket. That is the wrong number. It treats identity as a set of independent checkboxes and produces dashboards that flatter buyers while missing what audiences actually perceive: the interaction, the context, the intersection of age and skin tone and body size and disability in a single character. Vendors choose those coarse labels because they are easy to train and easy to demo. But the measurement a brand actually needs, whether its casting signals a lesbian couple or two friends, is exactly the measurement that coarse classifiers cannot make. There is also a quieter failure: most competitors sell private, unverifiable scores. Without public benchmarks, representation analytics becomes an audit you commission when you already know the answer. Alltold's willingness to publish quarterly industry numbers, with Adland.tv supplying the media repositories, makes their private reports credible in a way a closed score never is.
The takeaway for operators watching this space: generative AI has made representation measurable not optional. Once studios and platforms generate thousands of people per campaign, the only question is whether measurement is built into the pipeline or bolted on after the backlash. The companies that instrument it first, with methodology the measured communities actually endorse, will set the standard everyone else gets audited against.

