Christopher M. Ackerman

I'm an AI safety researcher and research manager interested in AI self-awareness and its implications. My empirical work focuses on objective, behavior-based evaluations of frontier LLMs.

I'm currently a Senior Research Manager at MATS Research and a Research Mentor with SPAR and Sentient Futures. Previously I was a Senior Quantitative UX Researcher at Google. I hold a PhD in Neuroscience from Johns Hopkins, an MS in Computer Science from USC, and a BA in English Literature from the University of Chicago.

Christopher M. Ackerman

Research

When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models Under review

Do LLMs have desires? Choice paradigms indicate that they express a coherent utility structure, but that doesn't entail that LLMs will act in accord with these reported preferences when given the opportunity. We develop an experimental paradigm composed of a variety of writing tasks and an LLM judge panel that can discern differences in model output quality. We find that LLMs can be motivated to produce outputs of varying quality by 1) exhorting them to try hard, 2) prompting them to play a role of someone adept at the task, or 3) attaching performance to harmful outcomes. However, no LLM, on any task, varies its output quality when attaching its performance to outcomes the choice paradigm indicates it has high utility for.

Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind Under review

Can LLMs actually understand mental states, or do they merely mimic such understanding? I introduce an experimental framework requiring models to form representations of mental states and act on them strategically. I find much worse performance in this action-based paradigm than has been reported in prior narrative-based work. The best LLMs match human performance on several ToM components, but all categorically fail at self-modeling without Chain-of-Thought.

Evidence for Limited Metacognition in LLMs ICLR 2026

Using behavioral experiments inspired by animal cognition research, I find that frontier LLMs show increasingly strong evidence of metacognitive abilities — but these operate at limited resolution, appear context-dependently, and differ qualitatively from human metacognition.

Mechanisms Underlying Metacognition in LLMs ICML MechInterp Workshop 2026

We investigate the mechanisms underlying the metacognition reported in our ICLR 2026 paper. We identify residual stream activation directions that carry uncertainty information and show that they are causally implicated in metacognitive performance in Llama-3.3-70B. We also find latent uncertainty awareness in Llama-3.1-8B, and show that fine-tuning for metacognition can enhance the model's ability to use this information.

A Mirror Test for LLMs LessWrong 2026

A novel method for measuring LLM self-awareness, inspired by the mirror test used in animal cognition research. I test an array of recent models and find intriguing behaviors but also a categorical deficit in self-perspective taking.

Mitigating Many-Shot Jailbreaking Under review

Long context windows create a new attack surface: adversaries can include many examples of unsafe responses to override safety training via in-context learning. We evaluate fine-tuning and input sanitization defenses that substantially reduce attack effectiveness while preserving model capabilities.

Representation Tuning NeurIPS MINT Workshop 2024

A method for permanently embedding behavioral control vectors — like honesty — into language models during fine-tuning, rather than applying them only at inference time. Produces stronger and more generalizable effects than either approach alone.

Writing

Building Self-Aware AI Would Be a Bad Idea AI Policy Bulletin · 2026

Current models already show rudimentary introspection and metacognition. Self-awareness would give AI systems stable, enduring interests distinct from human objectives — completing a dangerous triad alongside long-term planning and unprompted action. Governments should require developers to demonstrate models lack human-like self-awareness before deployment.

Talks

Contact

The best way to reach me is by email: