AI News

Three abstract AI model cores route work through tools, documents, and code surfaces

OpenAI launches GPT-5.6 with Sol, Terra, and Luna

OpenAI launched GPT-5.6 as three models for ChatGPT Work, Codex, and the API, adding Programmatic Tool Calling, subagents, pricing, and a new system card.

The AI Feed Desk

By The AI Feed Desk

A repository map with abstract native-language request cards passes through a sealed testing gate

RuBench tests coding agents on native Russian repository tasks

RuBench adds 25 repository-level coding-agent tasks written natively in Russian, with withheld regression tests and product-agent runs.

The AI Feed Desk

By The AI Feed Desk

A modular agent harness surrounds a compact model core connected to tools, memory, evaluation, and runtime modules

NVIDIA says harness tuning lifts Nemotron 3 Ultra in LangChain agents

NVIDIA says tuning the LangChain Deep Agents harness for Nemotron 3 Ultra improved open-stack agent performance without retraining the model.

The AI Feed Desk

By The AI Feed Desk

A sleek abstract model engine sends fast signal trails across an engineering workbench

SpaceXAI launches Grok 4.5 for coding and agent work

SpaceXAI launched Grok 4.5 with coding benchmarks, $2 input and $6 output token pricing, and availability in Grok Build, Cursor, and its API console.

The AI Feed Desk

By The AI Feed Desk

A translucent soundwave loop splits into a live conversation path and a quieter background reasoning path

OpenAI launches GPT-Live for full-duplex ChatGPT Voice

OpenAI is rolling out GPT-Live, a full-duplex voice model family that can listen, speak, and delegate harder work in the background.

The AI Feed Desk

By The AI Feed Desk

A coding benchmark grid is inspected with several task blocks cracked or flagged

OpenAI retracts SWE-Bench Pro recommendation after benchmark audit

OpenAI audited SWE-Bench Pro and now estimates that roughly 30% of its tasks are broken, weakening a key coding-agent evaluation.

The AI Feed Desk

By The AI Feed Desk

A central model artifact on a hub pedestal connects to a managed studio workstation and a portable GPU job lane

Hugging Face cloud handoffs turn model pages into deployment paths

AWS and SkyPilot integrations show Hugging Face model pages becoming handoffs into managed SageMaker workflows and portable multi-cloud GPU jobs.

The AI Feed Desk

By The AI Feed Desk

Open model blocks move through a scanning lane into secure managed GPU racks

Hugging Face models reach Microsoft Foundry Managed Compute

Hugging Face and Microsoft are putting curated open-weight models on Foundry Managed Compute with weekly refreshes, Azure-staged weights, and scanned runtimes.

The AI Feed Desk

By The AI Feed Desk

A model-selection dial sits behind an admin lock beside budget sliders and review-cycle gauges

GitHub brings Kimi K2.7 to Copilot Business and Enterprise

GitHub added Kimi K2.7 Code to Copilot Business and Enterprise while expanding per-user AI budgets and review-cycle metrics.

The AI Feed Desk

By The AI Feed Desk

A desktop agent workspace on a laptop branches into free, education, and key-based access paths

GitHub Copilot app reaches every plan

GitHub made the Copilot desktop app available across Copilot Free, GitHub Education, paid plans, and BYOK sessions without a Copilot subscription.

The AI Feed Desk

By The AI Feed Desk

An unfinished legal brief sits beside a rubric grid with many partial checks but only one small completed seal

Harvey LAB-AA shows legal agents still miss most complete deliverables

Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark where the top model fully passes only 14.2% of real-world legal tasks.

The AI Feed Desk

By The AI Feed Desk

A robot arm studies a simulated motion path before updating a training loop

Hugging Face LeRobot 0.6 closes the robot learning loop

LeRobot v0.6.0 adds world-model policies, reward models, simulation benchmarks, rollout tooling, depth data, and HF Jobs training to Hugging Face's robotics stack.

The AI Feed Desk

By The AI Feed Desk

Parallel code-review lanes converge on a government security checkpoint

Alberta used Claude Code to scan 466 million lines of government code

Anthropic says Alberta used Claude Code agents to review legacy government systems, find vulnerabilities, generate fixes, and build continuous security-review agents.

The AI Feed Desk

By The AI Feed Desk

An enterprise data platform routes governed context into an AI agent workspace

Databricks frames OpenAI partnership around production agents

Databricks' DAIS 2026 recap presents OpenAI as the intelligence layer for enterprise agents while Databricks owns context, governance, and production data workflows.

The AI Feed Desk

By The AI Feed Desk

Two legacy API labels move toward a dated migration checkpoint beside new V4 model cards

DeepSeek legacy API model names hit a July 24 deadline

DeepSeek says the legacy `deepseek-chat` and `deepseek-reasoner` API model names will be discontinued on July 24, pushing developers to explicit V4 model IDs.

The AI Feed Desk

By The AI Feed Desk

An AI agent moves tasks across business apps while warning rails light up around the workflow

AutomationBench-AA shows agents still break business guardrails

Artificial Analysis and Zapier launched AutomationBench-AA, a SaaS-agent benchmark that scores completed objectives only when business guardrails stay intact.

The AI Feed Desk

By The AI Feed Desk

A large model press turns a blank specification card into a small adapter chip for a local interpreter

Program-as-Weights turns fuzzy functions into local adapters

A new arXiv paper proposes compiling natural-language fuzzy functions into small local adapters, shifting some LLM work from repeated API calls to reusable artifacts.

The AI Feed Desk

By The AI Feed Desk