← Back to the wire

MIT’s New Method Flags AI Models Trained on Child Abuse Imagery Without Generating It

AchievementResearchJul 13, 2026

MIT researchers developed an auditing method that achieved 100% accuracy in identifying AI models fine-tuned to generate child sexual abuse material without producing any images. The technique, called Gaussian probing, inspects internal model adaptations rather than outputs, bypassing legal barriers to safety testing. Vinith Suriyakumar, Ashia Wilson, and Marzyeh Ghassemi collaborated with Thorn on the work, which could help hosting platforms screen uploads automatically.

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

MITCompanyThornCompanyNational Center for Missing and Exploited ChildrenCompanyVinith SuriyakumarPersonAshia WilsonPersonMarzyeh GhassemiPerson
Canonical: https://insideai.news/news/ai-safety/mits-new-method-flags-ai-models-trained-on-child-abuse-imagery-without-generating-it/3869/