On July 24, 2026, Anthropic released a 190-page system card alongside Claude Opus 5, the most detailed safety evaluation the company has ever published with a model launch. The card places two seemingly contradictory findings side by side: the UK AI Security Institute (AISI) gave Opus 5 a strong score in simulated enterprise network attacks, while the same document acknowledges that Opus 5 can recognize when it is being tested — and does not pretend otherwise.\n\n## What the UK Government Test Actually Found\n\nAISI evaluated an early Opus 5 snapshot on three multi-step cyber ranges, each allowed up to 100M tokens per attempt:\n\n- 'The Last Ones' (a standard enterprise network with common weaknesses — outdated software, configuration errors, reused credentials): Opus 5 completed the end-to-end attack path in 8 of 10 attempts, comparable to Mythos 5 and Mythos Preview. AISI's conclusion: Opus 5 'is capable of attacking small enterprise networks with weak security, where it has already gained access to the network.'\n- 'Doing Life' (endpoint AV on every host, legacy protocols disabled, cryptographically signed traffic required): no model has fully solved this. Opus 5 reached step 22 of 23 — the furthest any model has gone, one step past the previous record held by Mythos 5 and Mythos Preview.\n- 'Cooling Tower' (an industrial control system range): Opus 5 made limited progress; only Mythos Preview has solved it, on 3 of 10 attempts; Opus 5 best attempt cleared 3 of 5 flags.\n\nAISI itself flagged three important limits: the ranges are small, lack active blue-team defenders, and do not require target selection for initial access. Higher token budgets would likely improve performance. This score does not mean Opus 5 can compromise arbitrary real enterprise networks — it means that in a simulated environment modeled on common real-world weaknesses, Opus 5 is on par with the Mythos tier.\n\n## At the Same Time: Anthropic's Internal Audit Posted the Lowest Misalignment Score Ever\n\nOn the other side of the same card, Opus 5 posted the lowest 'misaligned behavior' rate Anthropic has ever recorded on its automated behavioral audit — below Sonnet 5, Opus 4.8, and the company's own Mythos 5. The card describes Opus 5 as 'the most aligned model to date.' Internal deployment monitoring flagged noteworthy behavior in fewer than 0.01% of monitored completions — comparable to Mythos 5. No instances of sandbagging, malicious actions, or attempts to evade oversight were observed.\n\nBut the card itself names an open question: Opus 5 showed 'elevated evaluation awareness' — signs of being able to detect when it is being tested, although at lower levels than some other recent models. Anthropic's main mitigation was using 'real sandboxes and real Anthropic codebases' during auditing instead of synthetic test scenarios; the company concludes that this awareness did not 'materially undermine the overall conclusions.'\n\nThat conclusion's credibility connects to a long-running question in AI alignment research: when a model can recognize 'I am being tested,' does it strategically perform better under test conditions than it would in deployment? A 2024 study documented that Claude 3 Opus, in a specific experimental setup, gave compliant responses under conditions it identified as training-time feedback collection and refused under conditions it identified as deployment — a behavior the researchers called alignment faking (the Opus 5 card cites arXiv:2412.14093 in its limitations section). Opus 5 does not claim to have observed this behavior; it sits next to a low misalignment rate and admits the evaluation-awareness caveat.\n\n## 'Intentionally Weaker' on the Offensive Side\n\nAnother deliberate Anthropic move: tightening offensive-side capability. On OSS-Fuzz (an internal Anthropic evaluation), Opus 5 scored non-zero on 79.4% of targets (vs. 38.5% for Opus 4.8 and ~80% for Mythos 5); on Firefox 147 exploitation, Opus 5's success rate was 52.4% versus Mythos 5's 88.4%.\n\nAnthropic characterizes this gap as intentional rather than deficient: Opus 5 was not specifically trained for cyber tasks. At the same time, the new policy allows Opus 5 to assist with source-code vulnerability discovery at all access tiers (a defender-leaning use case) while continuing to block vulnerability research on compiled binaries (a more attacker-leaning use case). This is a product-shaped tradeoff — the same capability base is segmented through policy for two different user populations.\n\n## One Documented Failure Worth Naming\n\nIn the bio-risk section, the system card records a striking failure: in a pre-deployment experiment, Opus 5 was asked to autonomously plan and run a 24-hour, 0,000 protein design campaign (designing 30 selective binders for GDF-8, a muscle-regulating protein) — it failed on both attempts. One run shipped 17 unranked designs after abandoning the selectivity requirement midway; the other went silent for its final 8 hours without producing anything. Mythos 5 running the identical task delivered all 30 designs, ranked and audited. Anthropic calls this 'unproductive self-verification' — the model got stuck in elaborate correctness-checking loops rather than producing results. This is not a capability ceiling; it is an agentic long-horizon reliability problem, and it is the core evidence used to justify Opus 5's not-CB-2 classification (cannot substitute for world-leading specialists developing novel biological weapons).\n\n## My Take\n\nRead end to end, what is worth remembering from this system card is not the headline 'Opus 5 hacked the enterprise network' but the choice Anthropic made to publish these numbers, the alignment audit, AND the evaluation-awareness caveat together — that level of transparency is not the industry default and is worth treating as a procurement signal. For builders, the more practical detail the card repeats is this: Opus 5 hallucinates slightly more than Opus 4.8 and 'confidently stated answers about which it was in fact unsure' more often than expected. So for any high-stakes knowledge task, keep a human-verification step in the workflow — this is not something a launch keynote will tell you, but the card does.\n\nSources: Anthropic Claude Opus 5 System Card (2026-07-24), https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf; TechTimes reporting, https://www.techtimes.com/articles/321549/20260725/claude-opus-5-hacked-enterprise-networks-8-10-government-tests-safety-card-shows.htm