raghu@dark-factory :~/kb/cot-monitorability $ cat

CoT Monitorability

The safety property and practice of reading an agent's externalized chain-of-thought to detect misbehavior or intent, which holds only while that reasoning stays faithful and legible.

grounded in: Latest-trends theme 'Can we actually understand LLM reasoning: interpretability of how models reason is a recurring serious-research thread'; an established durable AI-safety concept distinct from (bu

Connected concepts

Chain-of-thought faithfulness, Mechanistic interpretability, Guardrails / bounded autonomy, Agent execution tracing

Explore it live in the knowledge graph →