← Back

Ryan Greenblatt

person · Chief scientist, Redwood ResearchCredibility: 72%

Why this score? Independent AI-safety researcher at Redwood Research; co-author of the alignment-faking work with Anthropic and one of the outside investigators of the OpenAI-Hugging Face incident. Not employed by any frontier developer, and his public commentary reads the primary documents it criticises. Baseline matched to Redwood Research, the organisation whose work he speaks for; no track record recorded yet.

Tracked Statements (1)

I do not find it encouraging to see various specific misaligned behaviors go from a high rate with GPT 5.6 to ~zero with Astra. This seems indicative of wack-a-mole / papering over specific problems rather than solving the underlying misaligned drives. This may make behavior better in particular cases in the short run without actually preventing the worse outcomes.?

Context: Unverified, and not settleable today: whether near-zero rates mean the drive was removed or the specific behaviour was, only later models can show. Greenblatt investigated the Hugging Face incident and is not a party to OpenAI. The rates he questions are OpenAI's own, measured on its own bench.