AI Blog
· daily-digest · 5 min read

Google, Benchmarks, and Anthropic Under Pressure

Google is testing AI benchmarks in a double-blind setup, LAION is opening up a massive video dataset, and Anthropic is facing mounting legal pressure.

Inhaltsverzeichnis

Today’s focus is on three major topics that are currently shaking up the AI market: trust in benchmarks, open data as a raw material for research, and the legal headwinds facing Anthropic. On top of that, there are two Google updates showing that the company wants not just to keep up in AI research and safety, but to set standards.

If you want to know where AI development is headed right now, this mix of evaluation, data policy, and regulation is quite revealing. In short: less show, more system.

🔒 Google tests double-blind AI benchmarks

Google DeepMind is testing a frontier model in a double-blind evaluation for the first time – and that’s more important than it sounds at first glance. Using Confidential Space, the goal is to prevent Google from seeing the test questions or the evaluators from getting access to the model weights. In doing so, the company is addressing a real problem: many AI benchmarks are only as trustworthy as their setup. And unfortunately, that setup is often about as robust as an IKEA shelf after the third move.

In the pilot project with the Singapore AI Safety Institute, a Gemini Flash Lite is being used. The approach could help curb manipulation, leak risks, and “optimized” test conditions. For the industry, this could become a potential new gold standard for AI evaluation and confidential computing – especially at a time when benchmarks are increasingly becoming marketing tools rather than measurement instruments.

Source: The Decoder

🎥 LAION opens the Big Video Dataset for research

With the Big Video Dataset (BVD), LAION is providing one of the largest open video datasets for AI research. The scale is impressive: 80 million videos, around ten million hours of runtime, and 55 million clips with automatic descriptions. For video AI, this is a real raw-material boost, because good models need not only compute power, but above all large, diverse, and usable data.

The benchmark effect is also interesting: models trained on BVD are said to outperform the previous reference dataset InternVid by up to 2.1 percentage points. That may sound small, but in research it is often a fairly significant step. Legally, LAION is apparently relying on a Hamburg court ruling on data processing – a sign that open source, dataset research, and copyright remain closely intertwined. For anyone working on video models, this is: very relevant, very large, very data-hungry.

Source: The Decoder

🧪 DeepMind’s Co-Scientist becomes a research partner

Google DeepMind has evolved Co-Scientist from a hypothesis generator into a lab-integrated research system. According to the report, the Gemini-based multi-agent system is no longer just a suggestion engine for clever ideas, but is delivering experimentally validated results across three disciplines. From materials synthesis to the autonomous development of a medical AI architecture, this is no longer just an “interesting demo,” but real research support.

Why this matters: when AI not only summarizes texts but actively shapes the research process itself, it massively changes how work is done in labs. At that point, it’s no longer just about prompting, but about workflows, validation, and scientific responsibility. That’s the point where humans still matter – at least until the system starts thinking about the first coffee too.

Source: The Decoder

⚖️ Court rules Pentagon’s classification of Anthropic unlawful

A federal court in San Francisco has ruled that the Pentagon’s classification of Anthropic as a supply-chain risk was unlawful. The allegation is that the Department of Defense effectively blacklisted the company in retaliation for public criticism of the government’s AI policy. Formally, the classification remains in place for now because a parallel case is still pending in Washington – but the signal is clear.

This matters for Anthropic because the assessment is also tied to its planned IPO and the political environment. For the AI industry as a whole, the case is a reminder that regulation, public criticism, and corporate interests are now closely intertwined. Anyone aiming to grow in the AI sector must not only build models, but also be able to navigate legal waters. And against currents from multiple directions.

Source: The Decoder

Sony Music and Warner Chappell have filed a lawsuit against Anthropic in the US – over allegedly tens of thousands of copyrighted works. The plaintiffs are seeking up to $150,000 per work, plus additional penalties if identifiable copyright data was removed. In the worst case, we’re talking about damages in the billions. That’s not legally decided yet, of course, but the scale already shows that the copyright conflict around generative AI is long past being a side issue.

For you as an observer of the industry, this means: the legal framework for AI training remains one of the biggest sources of uncertainty. Especially tricky is that the question is not just “Is this allowed?” but also “What does it cost if it isn’t?”. For providers of generative AI, this is a clear sign that data provenance, licensing, and compliance are no longer side issues.

Source: The Verge

🛠️ Tool tip of the day

If you work with video or research data, it’s worth having a clean workflow for datasets, versioning, and experiments. Especially with open datasets like BVD or with multi-agent setups, reproducibility is gold. For such workflows, a tool for model and experiment management can quickly become useful – for example, take a look at #.

🧠 Research without a center: distributed learning in the background

A new arXiv preprint on “Decentralized Multitask Learning over Learned Task Graphs” shows where part of ML research is headed: away from centralized assumptions, toward distributed data and dynamically learned relationship structures. The paper investigates how multitask learning works when task relationships are unknown and must first be learned from distributed data.

This is especially interesting for real-world applications where data is distributed across different locations, devices, or organizations. That is exactly where many elegant lab models fail in practice. So the approach could become relevant for decentralized learning, edge environments, and networked systems. Bottom line: AI research is not only getting bigger, but also closer to actual systems. And that’s usually the point where it gets interesting.

Source: arXiv


Don’t want to miss any news? Subscribe to the newsletter


Weekly AI news highlights

No spam. No ads. Just the essentials — concisely summarized. Weekly in your inbox.