AI Blog
· daily-digest · 5 min read

Claude 5.5, AI Security and the Price of Open Models

Claude 5.5, critical vulnerabilities in coding agents, and new research on KV cache: today’s AI news with context and analysis.

Inhaltsverzeichnis

Today’s AI news revolves around three big themes: better models, more security, and the question of how “open” open models really are. Anthropic is rolling out the next model generation with Claude 5.5, while security researchers are simultaneously showing how quickly coding agents can become an entry point for attacks.

There are also exciting research papers on more efficient inference — exactly the kind of optimizations that ultimately determine whether AI products remain fast enough and affordable. In short: today is not just about new models, but also about the uncomfortable question of who is operating them safely.

🤖 Claude 5.5: Anthropic raises the bar on speed and price

Anthropic has introduced Claude Opus 5.5, the next generation of its Claude models. According to reports, the model is said to perform at the level of Claude Fable 5.1 on many tasks, while being around 40 percent cheaper to run. That is a very important point: with LLMs, it is no longer just quality that matters, but also the cost per request when deciding whether to use a model in a product.

The context is also interesting: with this release, Anthropic is signaling that competition is not only about being “better,” but above all about being “more efficient.” Faster responses, lower operating costs, and naturally sounding output are exactly the levers companies love — and whose bill nobody likes to look at at the end of the month. The announcement is also only the beginning: Sonnet 5.5 and Haiku 5.5 are expected to follow. For the market, that means the gap between top-tier models and production-ready deployment is shrinking, but pressure on prices is also increasing.

Source: The Decoder · Heise

🔐 Critical vulnerability in coding agents: autostart for malicious code

A new security vulnerability called Plugin4Shell affects several prominent coding agents and developer tools: Claude Code, OpenAI Codex, GitHub Copilot, and Gemini CLI. The core problem is unpleasantly simple: malicious code can be loaded and then executed automatically. Naturally, that is especially dangerous in tools that have access to files, shells, or repositories.

Why does this matter? Because coding agents in many teams are currently moving from experiment to productive assistant. The more privileges a tool has, the larger the attack surface becomes. According to the report, Anthropic and OpenAI have already patched the issue, but GitHub has not yet responded. For you as a user, that means taking updates seriously, keeping permissions tight, and not letting agents run everything blindly just because the output sounds friendly. AI can do a lot — but apparently it does not automatically look out for your security.

Source: Heise

🧠 H-Spec: speculative decoding without additional KV cache

The research paper H-Spec tackles one of the most important efficiency questions in large language models: how can inference be made faster without ruining output quality? Speculative decoding works roughly like this: a small draft model predicts several tokens ahead, which are then verified by the large model. This saves time because the expensive model does not have to generate every token individually.

H-Spec now takes this a step further and promises parallelization without additional KV cache on the draft side. This is technically relevant because KV cache is often one of the biggest drivers of memory usage and latency in LLM systems. If this approach proves itself in practice, it could be especially interesting for deployment scenarios with limited resources — in other words, everywhere model quality is not enough and responses also need to arrive quickly. Work like this is rarely glamorous, but it is exactly what turns demo AI into real product AI.

Source: arXiv

🧮 ValueDiff: smarter KV cache eviction

ValueDiff is also about inference optimization, this time focusing on KV cache eviction. Put simply: as a model processes longer context, the cache grows. At some point, that gets expensive. Existing methods often rely on so-called “attention sinks,” but these patterns can be less pronounced in modern models. Instead, ValueDiff uses the geometry of the value vectors to decide what should stay in the cache and what should be removed.

This is especially interesting for modern LLMs with QK normalization, gated attention, or logit softcapping — exactly the architectures becoming more common in current systems. The practical benefit is obvious: if cache management improves, latency and memory usage go down. For companies, that is not academic luxury, but real money. After all, anyone deploying LLMs at scale pays inference costs in actual cash, not in hypothetical currency.

Source: arXiv

🌍 Xiaomi builds a strong open model — with a catch

Xiaomi is making headlines with MiMo-V2.6-Pro: the model is said to be at the top among freely available AI models and was apparently trained surprisingly cheaply — with around 2.62 million US dollars for massive reinforcement learning. In today’s AI market, that is almost a bargain, especially considering how expensive large model training can otherwise be.

But: the story is apparently not entirely clean. According to the report, Anthropic accuses Xiaomi of having scraped training data from Claude. If confirmed, that would be another example of how closely powerful models and training data are intertwined these days. For the market, the lesson is: “open” is not automatically the same as “without a catch.” And for everyone relying on open models, the important question remains where the performance actually comes from — from clean training or from the large gray area in between.

Source: The Decoder

🛠️ Tool tip of the day

If you work with coding agents, it is worth taking a look at a tool for runtime monitoring and policy enforcement for AI workflows. Especially after Plugin4Shell, it is clear that agents should not only be useful, but also constrained. A tool that logs executions, enforces approvals, and secures sensitive actions can save you a lot of trouble in the long run. #


Don’t want to miss any news? Subscribe to the newsletter


Weekly AI news highlights

No spam. No ads. Just the essentials — concisely summarized. Weekly in your inbox.