AI News
Claude Science puts scientific AI inside a traceable workbench
Anthropic's Claude Science beta gives researchers a workbench with curated scientific skills, compute access, auditable artifacts, and reviewer agents.
By The AI Feed Desk
Hugging Face and Every Eval Ever make model-card scores more inspectable
Community Evals and Every Eval Ever now connect model-page benchmark scores to structured provenance records.
By The AI Feed Desk
Google gives coding agents an eval flywheel instead of another prompt tweak
Google's new quality-flywheel skill lets coding agents run structured agent evaluations with independent grading and production-trace loops.
By The AI Feed Desk
OpenAI's GeneBench-Pro makes biology benchmarks about judgment
GeneBench-Pro tests whether AI agents can handle ambiguous computational-biology analysis, not just clean benchmark questions.
By The AI Feed Desk
Anthropic restores Fable 5 and proposes a jailbreak severity framework
Fable 5 returns after US export controls were lifted, but the bigger change is Anthropic's push for a common way to score AI jailbreak risk.
By The AI Feed Desk
Claude Sonnet 5 turns Anthropic's default model into an agent model
Anthropic made Sonnet 5 the default for Free and Pro users while positioning it as a lower-cost agentic model close to Opus 4.8.
By The AI Feed Desk
Google UK ties its AI productivity case to worker training
Google's UK update pairs large economic-impact claims with an AI skills push aimed at closing the adoption gap.
By The AI Feed Desk
Claude reaches Microsoft Foundry with Azure governance and GB300 compute
Anthropic made Claude generally available in Microsoft Foundry, while NVIDIA framed the Azure deployment as a GB300 Blackwell Ultra agent platform.
By The AI Feed Desk
TraceLab turns real Codex and Claude Code sessions into serving data
The University of Washington TraceLab release studies coding agents as production workloads, with public traces across sessions, tool calls, tokens, cache behavior, and latency.
By The AI Feed Desk
Allen AI's DiScoFormer tests one transformer for density and score
The Hugging Face writeup frames DiScoFormer as a reusable estimator for density and score, with stronger high-dimensional results than kernel density estimation.
By The AI Feed Desk
AI data centers face a grid-connection bottleneck
A Works in Progress analysis argues that the limiting factor for AI buildouts is often the queue to connect new loads and generation to the electric grid.
By The AI Feed Desk
Google Finance turns portfolio tracking into an AI briefing workflow
Google Finance is adding portfolio ingestion, scheduled market briefings, and a new Android app as AI moves into recurring consumer finance tasks.
By The AI Feed Desk
Cognition's FrontierCode asks whether AI code would survive review
FrontierCode evaluates coding agents on mergeability, code quality, scope, tests, and maintainer judgment instead of only functional correctness.
By The AI Feed Desk
NVIDIA and Palantir put Nemotron open models inside air-gapped government AI
NVIDIA says Palantir is using Nemotron open models in isolated environments for U.S. government agencies and critical infrastructure operators.
By The AI Feed Desk
HP scales OpenAI Frontier from pilots into enterprise operating workflows
HP is expanding its OpenAI Frontier partnership after pilots in code, security, partner support, device operations, and employee workflows.
By The AI Feed Desk
OpenAI maps where EU jobs may grow, automate, or reorganize around AI
OpenAI's EU jobs framework splits AI labor impact into growth, automation, reorganization, and lower-change categories instead of treating exposure as one forecast.
By The AI Feed Desk
Hugging Face makes vLLM serving a one-command Jobs workflow
HF Jobs can now spin up a private OpenAI-compatible vLLM endpoint for tests, evals, and batch generation without provisioning servers or managing Kubernetes.
By The AI Feed Desk



















