AI Models

Anthropic Claude Opus 4.7 Boosts Coding Scores

Anthropic launched Claude Opus 4.7, scoring 64.3 percent on SWE-bench Pro, surpassing Opus 4.6's 53.4 percent and OpenAI's GPT-5.4 at 57.7 percent. The model triples image resolution to 2,576 pixels and reduces cybersecurity features during training with new blocks on risky requests. Pricing holds steady per token, but a new tokenizer raises effective costs.

Neura News

Neura News

Neura Market Editorial

April 16, 20263 min read
Anthropic Claude Opus 4.7 Boosts Coding Scores

Anthropic Claude Opus 4.7 Boosts Coding Scores

Anthropic released Claude Opus 4.7 as an upgrade to Opus 4.6. The model shows strong gains in autonomous coding. It achieved 64.3 percent on the SWE-bench Pro benchmark, compared to 53.4 percent for Opus 4.6. This score beats OpenAI's GPT-5.4 at 57.7 percent, though Claude Mythos Preview holds the top spot at 77.8 percent.

The company states that Opus 4.7 follows instructions with greater accuracy than before. Prompts designed for prior versions might yield different outcomes now. Opus 4.6 occasionally overlooked or loosely followed directions, but the new model takes them more literally.

Enhanced Image Processing

Opus 4.7 handles images up to 2,576 pixels along the longest edge. Anthropic calculates this at about 3.75 megapixels, over three times the capacity of previous Claude models. The change happens at the model level, so images process automatically at higher detail, using extra tokens. Users can reduce image size upfront if high resolution proves unnecessary.

Anthropic points to benefits for agents that analyze screenshots or diagrams. On the OfficeQA Pro benchmark for document reasoning, accuracy reached 80.6 percent, a jump from 57.1 percent with Opus 4.6. Results also improved in biomolecular reasoning and the ScreenSpot-Pro visual navigation test.

Limits on Cybersecurity Features

Anthropic took steps to curb specific cybersecurity abilities in Opus 4.7. During training, the company worked to lower these capabilities on purpose. Safeguards now spot and stop requests linked to prohibited or dangerous cyber activities.

This approach connects to Project Glasswing, where Anthropic examined AI risks and upsides in cybersecurity. The firm chose to limit Mythos Preview's rollout and test protections first on models like Opus 4.7.

Security experts interested in penetration testing or red-teaming can join the Cyber Verification Program.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Fewer Hallucinations and Alignment Notes

The system card separates factual hallucinations, such as false world claims or invented quotes, from input hallucinations, where the model assumes unavailable tools exist.

Opus 4.7 matches or exceeds Opus 4.6 on factual hallucinations across four tests, though it trails Mythos Preview. The difference stems from Mythos Preview's edge on rare facts, not more errors in Opus 4.7.

For input hallucinations, Opus 4.7 records the lowest rate among tested models for missing tools. It nears Mythos Preview when context lacks and outperforms older versions. Tests focused on Opus 4.6 flaws, which affected those scores. On fabricated facts, it ties Opus 4.6 and lags Mythos Preview. Under duress from prompts to contradict itself, Opus 4.7 shows more honesty than Opus 4.6 but less resolve than Mythos Preview.

Safety remains close to Opus 4.6 overall, with low deception, sycophancy, and misuse cooperation. It resists prompt injections better. One ongoing problem: it declines 33 percent of simulated AI safety research tasks, down sharply from 88 percent for Opus 4.6.

Pricing and New Options

Input tokens cost $5 per million, output $25 per million, unchanged from before. A fresh tokenizer assigns up to 1.35 times more tokens to identical text. Higher effort prompts produce longer outputs, so real per-request costs can climb.

A new "xhigh" effort level fits between "high" and "max." Claude Code adds "/ultrareview" for code checks and broadens "Auto Mode" for Max users, allowing independent decisions. Access comes via Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. Check the migration guide for Opus 4.7 details.

Anthropic, founded in 2021 by former OpenAI staff, focuses on safe AI systems. Its Claude series emphasizes reliability and alignment in practical uses.

Related on Neura Market

More from Neura News