On May 6, OpenAI officially pushed GPT-5.5 Instant as the default model for ChatGPT, replacing the previous GPT-5.3 Instant. This is another important upgrade from the company on LLM reliability.
Core improvement: sharp drop in hallucination rate
OpenAI's emphasis this time is on reliability improvements in high-risk domains like law, medicine, and finance. Compared to its predecessor, GPT-5.5 Instant significantly reduces the frequency of incorrect answers while maintaining low latency. This is significant for professional scenarios requiring accurate information.
Performance data is impressive
On the AIME 2025 math test, GPT-5.5 Instant scored 81.2, far above GPT-5.3's 65.4. On the MMMU-Pro multimodal reasoning benchmark, the score also rose from 69.2 to 76. This set of data shows OpenAI has achieved a comprehensive capability boost while maintaining response speed.
Context management becomes the biggest highlight
The new model supports retrieving context across conversations, files, Gmail, and other sources, producing more personalized answers. This feature is currently available to Plus and Pro users on web, with mobile launching soon. Free users and enterprise users are expected to gain access in the coming weeks.
Additionally, ChatGPT will display memory sources across all models, helping users track the origin of answers, enhancing transparency.
Impact on developers
When accessing via API, developers can use "chat-latest" to point to GPT-5.5, while GPT-5.3 will be removed from paid tiers in three months.
Commentary
GPT-5.5 Instant reflects a clear direction: as LLM capability approaches its ceiling, improving reliability and user experience is becoming the new competitive dimension. Hallucination was once one of the most frequent angles outsiders used to question LLMs, and this targeted optimization may redefine the industry's standard for "trustworthy AI." OpenAI isn't blindly chasing benchmark numbers this time, but landing improvements in real use scenarios — a strategic shift worth attention.