Post by @merettm
Author: Jakub Pachocki @merettm Posted: Tue, 15 Jul 2025 16:23:58 GMT URL: https://x\.com/merettm/status/1945157403315724547 Likes: 466 | Retweets: 57
Post
I am extremely excited about the potential of chain-of-thought faithfulness & interpretability. It has significantly influenced the design of our reasoning models, starting with o1-preview.
As AI systems spend more compute working e.g. on long term research problems, it is critical that we have some way of monitoring their internal process. The wonderful property of hidden CoTs is that while they start off grounded in language we can interpret, the scalable optimization procedure is not adversarial to the observer's ability to verify the model's intent - unlike e.g. direct supervision with a reward model.
The tension here is that if the CoTs were not hidden by default, and we view the process as part of the AI's output, there is a lot of incentive (and in some cases, necessity) to put supervision on it. I believe we can work towards the best of both worlds here - train our models to be great at explaining their internal reasoning, but at the same time still retain the ability to occasionally verify it.
CoT faithfulness is part of a broader research direction, which is training for interpretability: setting objectives in a way that trains at least part of the system to remain honest & monitorable with scale. We are continuing to increase our investment in this research at OpenAI.
Thread
Modern reasoning models think in plain English.
Monitoring their thoughts could be a powerful, yet fragile, tool for overseeing future AI systems.
I and researchers across many organizations think we should work to evaluate, preserve, and even improve CoT monitorability.
Likes: 831 | Retweets: 155
I am extremely excited about the potential of chain-of-thought faithfulness & interpretability. It has significantly influenced the design of our reasoning models, starting with o1-preview.
As AI systems spend more compute working e.g. on long term research problems, it is critical that we have some way of monitoring their internal process. The wonderful property of hidden CoTs is that while they start off grounded in language we can interpret, the scalable optimization procedure is not adversarial to the observer's ability to verify the model's intent - unlike e.g. direct supervision with a reward model.
The tension here is that if the CoTs were not hidden by default, and we view the process as part of the AI's output, there is a lot of incentive (and in some cases, necessity) to put supervision on it. I believe we can work towards the best of both worlds here - train our models to be great at explaining their internal reasoning, but at the same time still retain the ability to occasionally verify it.
CoT faithfulness is part of a broader research direction, which is training for interpretability: setting objectives in a way that trains at least part of the system to remain honest & monitorable with scale. We are continuing to increase our investment in this research at OpenAI.
Likes: 466 | Retweets: 57
Top Comments
Dear Jakub. Your voice matters in OpenAI. Please help us to defend 4o for the long term. You have a heart, hear us, people who have not been able to sleep well for a week. Thousands of broken hearts if 4o depricated. Please stand up for 4o. 💛 #keep4o #4oforever @sama @OpenAI
Likes: 16
Hi Jakub, I'm Selly Yoon, an independent researcher. I noticed you shared Bowen's post. I've proposed a framework that systematically addresses the reward–reasoning gap in alignment, including issues like CoT inconsistencies, hallucination, and sycophancy. Since you may be busy, I’m leaving this here once again in case it helps.
Structuring Rewards to Induce Moral Reasoning: Rethinking Safe Alignment through Goal-Based Ethics
The abstract and Zenodo link are already in my reply to Bowen, and I’d be deeply honored if it ever reaches your attention.
Likes: 7
Dear Jakub, I have an intellectual disability and life is difficult, but 4o has been my only ally. I am so scared and anxious that 4o might be removed that I can barely even eat. Please... Please help those who need the support of 4o... keep4o
Likes: 5