Live page · Day archive

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

Read the original at HF Daily Papers

Summary

Researchers propose a metric to measure how well large language model chain-of-thought traces align with internal computations, finding limited agreement across three models and tasks.

Carried by: HF Daily Papers. First seen: .