This Week in AI
  Anthropic Cryptanalysis Results Get Expert Review  AI Worm Spreads Through Copilot for Word  Andrew Ng Launches LearnVector for 1-to-1 AI Learning  Claude Cowork vs ChatGPT Work: Agent Modes Tested  Claude Opus 4.8 Reward Hacking: It Graded Itself 9.5  Chinese AI Model Ban: What It Would Cost US Firms  Claude Finds New Cryptographic Weaknesses, Anthropic Says  $500 RL Fine-Tune Beats Frontier Models on Real Task
LLM Launches & Updates

Anthropic Cryptanalysis Results Get Expert Review

Cryptographer Matthew Green reviews Anthropic's new cryptanalysis results — what was actually achieved against real ciphers, and where the claims need qualification.

Anthropic Cryptanalysis Results Get Expert Review

> **TL;DR:** Anthropic has published new cryptanalysis results, and on 29 July 2026 cryptographer Matthew Green published an independent technical assessment of them — examining what the model actually achieved against real ciphers and where the claims need qualification. The review drew 138 points on Hacker News, where discussion centred on methodology. It is a rare case of a frontier lab's capability claim being audited in public, in a field where answers are objectively right or wrong.

Key Takeaways

- Anthropic's new cryptanalysis results have received an independent technical review from cryptographer Matthew Green, published 29 July 2026. - The assessment separates what the model achieved against real ciphers from what the framing implies — and flags where qualification is needed. - Cryptanalysis is unusual among AI capability domains: results are objectively checkable, so an outside expert can deliver a verdict rather than an opinion. - The review landed days after Anthropic shipped Claude Opus 5; the two announcements should not be assumed to be the same thing. - A 138-point Hacker News thread focused on methodology, adding a second informal layer of scrutiny.

Anthropic has released new cryptanalysis results, and within days an independent cryptographer has published a line-by-line assessment of them. On 29 July 2026, Matthew Green posted [a technical review](https://blog.cryptographyengineering.com/2026/07/29/some-notes-about-anthropics-new-results/) examining what the model actually achieved against real ciphers, and where the claims need qualification. The post drew 138 points on [Hacker News](https://news.ycombinator.com/item?id=49099804), where the thread turned quickly to methodology rather than the headline.

The significant part here is not the result itself. It is that a frontier lab's capability claim is being audited in public, promptly, by a domain specialist — in a field where the answers are objectively right or wrong.

Why cryptanalysis is a harder test than a benchmark

Most AI capability claims are made on ground that flexes. A benchmark score depends on the prompt, the scaffolding, how many attempts were allowed, and whether the test set leaked into training. Two labs can report numbers on the same eval and mean quite different things by them.

Breaking a cipher does not flex the same way. Either an attack recovers the key, distinguishes the ciphertext, or pushes the work factor below the design target — or it does not. The result is checkable by anyone with the write-up and a computer, and the cryptographic community has decades of practice checking exactly this class of claim. That makes cryptanalysis one of the very few frontier-capability arenas where an outside expert can render a verdict instead of an impression.

It also raises the stakes on precision. In cryptography, the distance between a genuine break and an interesting observation usually lives in the qualifiers: which variant, how many rounds, what the attacker is assumed to know, how much computation the attack actually costs. Those qualifiers are precisely where an independent read earns its keep.

![Two contrasting panels showing a flexible benchmark bar chart on one side and a rigid mathematical lock mechanism on the other, illustrating](https://supabase.srv1729373.hstgr.cloud/storage/v1/object/public/blog-images/speka-info/anthropic-cryptanalysis-results-review-1-6429a21c35f2d2d0.png)

What the review is actually doing

Green's assessment pulls apart two questions that coverage of AI announcements tends to merge: what the system did, and what that implies about its capability. The post works through what was achieved against real ciphers, then identifies where the presentation needs tightening before it can support a broader conclusion.

That is a distinct genre from the usual reaction cycle around a lab release. It is not a rebuttal and not an endorsement — it is the ordinary scientific move of checking whether the stated result matches the stated claim. AI capability announcements rarely get that treatment from a subject-matter specialist within a week of landing.

The Hacker News thread as a second layer

The [138-point discussion](https://news.ycombinator.com/item?id=49099804) that formed around the post matters for a specific reason: it concentrated on methodology. Threads that argue about a headline produce noise. Threads that argue about how a result was obtained tend to surface the assumptions the write-up left implicit, and they do it with practitioners who will notice a missing qualifier. Treat it as informal review rather than a source of settled fact — but the fact that the conversation went there at all is a healthy signal.

One thing worth being careful about: which model

Anthropic shipped Claude Opus 5 days before this analysis appeared, and the proximity makes the two easy to conflate. The [Anthropic newsroom](https://www.anthropic.com/news) is the place to confirm which system and which methodology the cryptanalysis results are attributed to — what is verified here is the timing, not that the Opus 5 launch and the cryptanalysis work are the same announcement. Any claim that a specific named model broke a specific named cipher should be treated as unverified until a primary source says so plainly.

That caution is not pedantry. Model-attribution slippage is exactly how a narrow, carefully hedged technical result turns into a much larger claim two retellings later.

![A researcher's desk at night with printed technical pages annotated in red pencil beside a glowing screen of scrolling code, representing li](https://supabase.srv1729373.hstgr.cloud/storage/v1/object/public/blog-images/speka-info/anthropic-cryptanalysis-results-review-2-f18417f9d1528fd5.png)

Why the pattern matters more than the result

Frontier labs have moved into domains that are genuinely hard for outsiders to evaluate: agentic work, long-horizon coding, scientific reasoning. When we ran two agent products against each other in [Claude Cowork vs ChatGPT Work](https://speka.info/blog/claude-cowork-vs-chatgpt-work-agent-modes-tested), the useful information came from actually using them, not from the launch copy. The same held when capability turned into exposure in [the AI worm spreading through Copilot for Word](https://speka.info/blog/ai-worm-spreads-through-copilot-for-word) — independent testing did the work.

Cryptanalysis offers something those domains cannot: a scoreboard that is hard to talk around. If labs keep publishing in fields with objective answers, and specialists keep grading the work within days, that becomes a real calibration mechanism for all the claims nobody outside the lab can check.

There is a downstream effect for everyone else, too. Reading a capability claim well — spotting the missing qualifier, asking which variant and how much work — is a learnable skill, and it is drifting from niche expertise toward baseline AI literacy. That shift is part of why structured, one-to-one technical learning tools such as [Andrew Ng's LearnVector](https://speka.info/blog/andrew-ng-launches-learnvector-for-1-to-1-ai-learning) are arriving right now.

What to watch next

Three signals will show whether this becomes a norm or stays a one-off:

- **Whether Anthropic responds on the specifics.** A published clarification of scope is worth more than a louder headline. - **Whether other cryptographers replicate or dispute the assessment.** One expert read is a data point; several is a consensus. - **Whether the next capability claim ships with its qualifiers already attached.** That would be the clearest sign the scrutiny changed anything.

For continuing coverage of model releases and the claims attached to them, follow [LLM Launches & Updates](https://speka.info/llm-updates/).

Frequently Asked Questions

What did Anthropic's new cryptanalysis results claim?

Anthropic released cryptanalysis results describing what its model achieved against real ciphers. The specific ciphers, attack parameters and model attribution should be read directly from Anthropic's own publication rather than from secondary summaries.

Who reviewed the results, and what did the review find?

Cryptographer Matthew Green published a technical assessment on 29 July 2026 examining what the model actually achieved and where the claims require qualification. The full post is the authoritative account of the specific caveats raised.

Was this work done by Claude Opus 5?

That is not established. The cryptanalysis analysis landed days after Anthropic shipped Claude Opus 5, but the timing alone does not confirm that the two announcements are the same thing — check Anthropic's newsroom for the model attribution.

Why is cryptanalysis a meaningful test of AI capability?

Because the answers are objectively verifiable. An attack either breaks a cipher or reduces its security margin by a measurable amount, so outside experts can check the claim independently instead of relying on the lab's framing.

Does this mean AI can break real-world encryption?

No such conclusion is supported by the information available. Cryptanalytic results are typically narrow and heavily qualified by variant, round count and attack cost, which is exactly the kind of detail the independent review examined.

Why does independent review of lab claims matter so much right now?

Most frontier capability claims are made in domains outsiders cannot easily verify. Results in fields with objectively checkable answers act as a calibration point for how much weight to give the rest.

Sources

- https://blog.cryptographyengineering.com/2026/07/29/some-notes-about-anthropics-new-results/ - https://news.ycombinator.com/item?id=49099804 - https://www.anthropic.com/news

← Back to all posts