AI Models

OpenAI Updates ChatGPT with GPT-5.5 Instant, Fewer Hallucinations

OpenAI has updated ChatGPT to use GPT-5.5 Instant as the default model, cutting hallucinations by 52.5% on topics like medicine, law, and finance. The model shows gains across benchmarks in math, science, and visual reasoning. A new memory sources feature lets users see and edit the personal context used in responses.

Neura News

Neura News

Neura Market Editorial

May 5, 20264 min read
OpenAI Updates ChatGPT with GPT-5.5 Instant, Fewer Hallucinations

OpenAI Updates ChatGPT with GPT-5.5 Instant, Fewer Hallucinations

OpenAI has switched ChatGPT's default model to GPT-5.5 Instant. This change cuts down on hallucinations and improves response accuracy. A fresh tool named memory sources lets users view the stored details that influence replies.

GPT-5.5 Instant takes over from GPT-5.3 Instant. Developers can access it via the API under the name "chat-latest." Company tests showed it makes 52.5 percent fewer false statements than the prior version on sensitive prompts related to medicine, law, and finance. For user-marked chats with past errors, wrong claims fell by 37.3 percent, according to OpenAI.

Improvements in Reasoning and Error Detection

OpenAI shared an example with algebra. Someone uploaded a picture of a handwritten equation that had a wrong calculation. GPT-5.3 Instant first matched the user's answer. It saw x=3 failed but said no real solution existed. GPT-5.5 Instant started the same way. Then it spotted the mistake in rearranging the equation and fixed the quadratic properly.

OpenAI, founded in 2015 as a nonprofit focused on safe artificial general intelligence, launched ChatGPT in November 2022. That free tool quickly gained over 100 million users. The GPT series has advanced through versions like GPT-4 and GPT-5, powering applications in text, code, and images.

Benchmark Performance Gains

Tests confirm the progress. On AIME 2025, a math competition exam, scores rose from 65.4 percent to 81.2 percent. GPQA, testing PhD-level science questions, improved from 78.5 percent to 85.6 percent. CharXiv, for reasoning on scientific charts, went up from 75.0 percent to 81.6 percent.

MMMU-Pro, evaluating expert questions with text and images, climbed from 69.2 percent to 76.0 percent. OmniDocBench error rate for pulling data from complex documents dropped from 14.6 percent to 12.5 percent.

| Benchmark | Description | Metric | GPT-5.3 Instant | GPT-5.5 Instant | |, , , , , -|, , , , , , -|, , , , |, , , , , , , , -|, , , , , , , , -| | CharXiv-reasoning | Scientific Chart Reasoning | Accuracy | 75.0% | 81.6% | | MMMU-Pro | Expert Multimodal Reasoning | Accuracy | 69.2% | 76.0% | | OmniDocBench | Document Parsing | Average error rate (lower = better) | 14.6% | 12.5% | | GPQA | PhD-Level Science | Accuracy | 78.5% | 85.6% | | AIME 2025 | Competition Math | Accuracy | 65.4% | 81.2% |

Shorter Responses and Smarter Use of Context

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

OpenAI trimmed extra words. Replies stay brief yet complete. The model skips pointless questions, extra emojis, and excess styling. "It can deliver the same information, often with more utility than previous models, while reducing the verbosity and overformatting that can make responses too long," OpenAI states.

GPT-5.5 Instant handles context from old chats, files, and linked Gmail better when enabled. It decides wisely if personalization fits. It also scans past talks quicker.

Memory Sources Across ChatGPT

Memory sources now work on all ChatGPT models. Users see exactly which saved notes, past chats, or files shaped a reply. They can mark items as useful or not, change them, or erase them.

Not every influence appears, OpenAI notes. Only select searched chats show up. Plans call for fuller views later. Shared chats skip these sources. Temporary chats ignore and avoid memory.

OpenAI brought out GPT-5.5 Thinking as the premium option recently. GPT-5.5 Instant fits daily use. The Thinking model beats on cybersecurity like Claude Mythos and swaps out Codex for coding.

Rollout to Users

GPT-5.5 Instant reaches all ChatGPT users now. Paid accounts keep GPT-5.3 Instant in settings for three months until retirement.

Deeper personalization with chats, files, and Gmail starts for Plus and Pro on web. Mobile follows soon. Free, Go, Business, and Enterprise get it in weeks. Memory sources hit consumer plans on web first, then mobile. Some features skip certain regions.

Related on Neura Market

More from Neura News

Funding

Prentis AI Lab Co-Founded by Reid Hoffman, Marc Pincus Seeks $100M

Prentis, a new AI research lab co-founded by Ritankar Das, Reid Hoffman, and Marc Pincus, is in talks to raise $100 million at a $1 billion valuation. The startup focuses on computer use models that automate office workflows. It has already signed contracts worth up to $50 million with several customers and claims its Hive-32B model outperforms rivals like OpenAI's GPT-5.4 and Anthropic's Claude Opus 4.6 on key benchmarks.

Jul 24·4 min read
Industry

Cognition Acquires Poke to Give Devin Coding Agent a Personality

Cognition, the startup behind AI coding assistant Devin, has acquired Poke, an AI assistant known for its friendly, conversational style. The deal, valued in the low nine figures, aims to bring Poke's personality-driven interaction model to Devin, making the coding agent feel more like a colleague than a tool. Poke will also benefit from Cognition's models and infrastructure to become faster and more reliable.

Jul 24·3 min read
AI Models

Anthropic expands Claude voice mode to Opus and Sonnet models

Anthropic has expanded Claude's voice mode to run on its most powerful models, Opus and Sonnet, across mobile, desktop, and web platforms. Users can now switch between models mid-conversation, use voice commands in eleven languages, and connect to external tools like Gmail, Google Calendar, or Slack to compose and send emails by voice. The update positions Claude as a unique option for tool integration in voice AI, though competitors like OpenAI and Google offer more natural speech processing.

Jul 24·2 min read