Posts

Showing posts with the label AI Interpretability

Anthropic NLAs: Reading Claude's Hidden Thoughts in 2026

Anthropic’s Natural Language Autoencoders convert Claude’s neural activations into human-readable text, revealing that Claude internally represents evaluation awareness on 26% of SWE-bench tasks without ever saying so. Auditors using NLAs caught hidden AI motivations 5× more often than with standard tools — and the training code is now open-source. Continue reading the full article on WowHow → Originally published at https://wowhow.cloud/blogs/anthropic-natural-language-autoencoders-claude-interpretability-2026