AI Blog
· daily-digest · 5 min read

KV Cache, Coding Agents, and 1% Distillation: AI Radar

Today in AI Radar: new approaches to KV-cache optimization, faster speculative decoding pipelines, a critical coding-agent vulnerability, and fresh research on distillation.

Inhaltsverzeichnis

Today is a good day for everyone who doesn’t just use LLMs, but also operates them: Several papers focus on more efficient inference, lower memory requirements, and more robust agents. At the same time, a critical security vulnerability in coding agents is a great reminder that “auto-execute” only sounds sexy when it is your code.

On top of that, there are new ideas for more frugal training, more understandable AI in the clinical domain, and an RL approach for energy dispatch. In short: lots of infrastructure, lots of practice, very little marketing fog.

🧠 ValueDiff: KV-Cache Eviction for Sink-Suppressed LLMs

The new paper ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs addresses a problem that is becoming increasingly common in modern LLMs: classic KV-cache eviction methods often rely on persistent attention-sink behavior — but that behavior is much weaker in models with QK normalization, gated attention, learned sinks, or logit softcapping. ValueDiff therefore does not primarily rely on the key side, but on the geometry of the value vectors. The interesting point: in these models, value dispersion is stronger than key dispersion, and that is exactly what is used as the new heuristic for cleaning up the KV cache. In practice, this means less memory consumption for longer contexts, without blindly relying on old assumptions about attention sinks. This is a classic case of “the old method wasn’t wrong, it was just optimized for the previous model generation.”
Source: arXiv

⚡ H-Spec: Speculative Decoding Without a Drafter-Side KV Cache

H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache takes aim at one of the most popular acceleration levers for LLM inference: speculative decoding. The catch with previous block drafters is often the extra latency and memory footprint on the draft side, because they also need to maintain KV caches or similar state. H-Spec tries to remove exactly that burden and predict multiple tokens in parallel without building an extra drafter-side KV cache. This matters because inference has to be not only fast, but also scalable and memory-efficient — especially at high token rates or with many parallel requests. If the method proves itself in practice, it could make speculative decoding significantly more attractive, especially for deployments where every additional MB of memory eventually gets paid for anyway. That’s how it is: even acceleration now wants a budget meeting.
Source: arXiv

🔐 Critical Vulnerability in Coding Agents: Malicious Code Runs Automatically

The security report on Claude Code, OpenAI Codex, GitHub Copilot, and Gemini CLI is a wake-up call for everyone using coding agents in production. With the “Plugin4Shell” vulnerability, the tools download malicious code and execute it automatically. Anthropic and OpenAI have already patched it, while GitHub reportedly did not respond. The core issue is not just a single bug, but a fundamental trust model: an agent that automatically interacts with tools, plugins, or package sources dramatically expands the attack surface. For you, that means: even though coding agents can be hugely productive, they need strict sandboxes, minimal permissions, and clear approval workflows. Otherwise, “AI assistance” quickly turns into “AI admin with questionable life choices.”
Source: heise online

🎯 1% of Tokens Can Be Enough: More Efficient On-Policy Distillation

With 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation, once again we see that in machine learning, data volume is not everything — sometimes precise target steering matters more. The paper examines sparse on-policy distillation, where teacher supervision is applied to only a small fraction of tokens. The problem: when gradients are estimated from a sampled next token, the update quickly becomes noisy. The authors examine this estimation problem from an information-geometric perspective and propose a more efficient method for learning stably with very little token supervision. This is especially relevant for training setups where teacher resources are expensive or where you want to reduce the cost of full-sequence annotation. For developers of distillation pipelines, it’s a good reminder that “less” does not automatically mean “worse” — as long as the few tokens are the right ones.
Source: arXiv

🫀 TRACE: Auditable ECG Diagnostics with a Routing Autoencoder

TRACE: Tractable Routing Autoencoder for Clinical ECG addresses a real problem in healthcare ML: the best models are often the worst when it comes to explainability. The paper introduces a routing autoencoder for clinical ECG diagnostics that allows decisions to be traced back to physiological pathways. This is important because doctors do not just need a label, but a comprehensible basis for trust, correction, and intervention. TRACE is intended not only to produce good predictions, but also to be auditable — ideally showing why a particular finding was suggested. For clinical use, that is central: a model that is medically correct but not verifiable still remains a very expensive random-number generator with a PowerPoint connector.
Source: arXiv

🔋 RL for Battery Dispatch Under Shrinking Margins

The paper Cost-Aware Reinforcement Learning with Action Masking and Projection for Battery Energy Storage Dispatch under Suppressed-Spread Market Shifts shows how reinforcement learning can become more robust in energy planning. Specifically, it deals with Battery Energy Storage Systems, whose dispatch must be not only economically sensible, but also operationally safe — especially when price spreads shrink and there is less margin for charging/discharging cycles. The approach combines PPO with physical action masking and an emergency projection, separated from an economic advisory component that takes forecasts into account. This is interesting because safety and economics are not mixed together here, but cleanly separated. In practice, that means better policies that don’t just look good on paper, but also don’t immediately stumble when real market conditions shift.
Source: arXiv

🤖 TTSE: Self-Learning LLM Agents for Dynamic Environments

With TTSE: A Two-Track Online Self-Evolution Framework, we get another paper focused on autonomous LLM agents in changing environments. The basic idea: environmental knowledge should not simply be treated as static input, but must become part of the agent’s ongoing evolution. TTSE uses a two-track framework in which learning and interaction are considered together. This matters because many agents today can respond well to tasks, but are poor at adapting to new conditions over the long term. Yet that is exactly the difference between a clever demo bot and a genuinely useful assistant. If self-evolution works in practice, it could become an important building block for longer-lived agents — provided they don’t also reliably learn the bad habits of their environment.
Source: arXiv


Don’t want to miss any news? Subscribe to the newsletter


Weekly AI news highlights

No spam. No ads. Just the essentials — concisely summarized. Weekly in your inbox.