Large Language Models News

Developer work lanes compare coding, agent tasks, reasoning effort, and billing controls for Gemini 3.6 Flash in Copilot

Gemini 3.6 Flash rolls into GitHub Copilot

GitHub says Google's Gemini 3.6 Flash is rolling out in Copilot with configurable reasoning effort and parallel tool use.

The AI Feed Desk

By The AI Feed Desk

A security operations console rotates token keys into a vault beside a dataset processing pipeline

Hugging Face says an autonomous agent breached production infrastructure

Hugging Face disclosed a July 2026 production incident it says was driven by an autonomous AI agent system and recommends token rotation.

The AI Feed Desk

By The AI Feed Desk

An open model engine sits on a lab table with layered context bands and a release-calendar marker

Moonshot launches Kimi K3 as an open 2.8T model

Moonshot's Kimi K3 pairs a 2.8T-parameter model, 1M-token context, API access, and a promised July 27 weight release.

The AI Feed Desk

By The AI Feed Desk

Hot MoE experts move between SRAM, HBM, and shared DRAM tiers on a chiplet package

HCRMap targets hot-expert bottlenecks in MoE inference

A July 13 paper proposes HCRMap, a pressure-aware residency system for hot MoE experts across 3.5D chiplet memory tiers.

The AI Feed Desk

By The AI Feed Desk

A central managed agent node connects to four external tool and background task endpoints

Google adds background tasks and remote MCP to Gemini Managed Agents

Google expanded Managed Agents in the Gemini API with asynchronous background execution, remote MCP servers, custom functions, and credential refresh.

The AI Feed Desk

By The AI Feed Desk

A productivity workspace connects three app tiles to a central approved model marker

GPT-5.6 becomes Microsoft 365 Copilot's preferred model

OpenAI says GPT-5.6 is now the preferred model in Microsoft 365 Copilot across Word, Excel, PowerPoint, Chat, and Cowork.

The AI Feed Desk

By The AI Feed Desk

An abstract model graph passes through a glass optimization lens into high-speed inference lanes

Hugging Face and vLLM bring native-speed serving to Transformers model definitions

Hugging Face and vLLM introduced a backend that can run compatible Transformers model definitions at native vLLM speed through runtime graph analysis and rewrites.

The AI Feed Desk

By The AI Feed Desk

A model selector dial routes three model chips to laptop, terminal, cloud agent, and mobile coding surfaces

GitHub rolls GPT-5.6 into Copilot across IDEs, CLI, agents, and mobile

GitHub is adding OpenAI's GPT-5.6 Sol, Terra, and Luna to Copilot across major developer surfaces, with enterprise and business access disabled by default.

The AI Feed Desk

By The AI Feed Desk

A translucent model core connects long-context streams to image, video, code, and document stations

Meta launches Muse Spark 1.1 and opens its Model API preview

Meta launched Muse Spark 1.1 with a public Model API preview, positioning the model for coding, tool use, multimodal agent workflows, and long-context work.

The AI Feed Desk

By The AI Feed Desk

Three abstract AI model cores route work through tools, documents, and code surfaces

OpenAI launches GPT-5.6 with Sol, Terra, and Luna

OpenAI launched GPT-5.6 as three models for ChatGPT Work, Codex, and the API, adding Programmatic Tool Calling, subagents, pricing, and a new system card.

The AI Feed Desk

By The AI Feed Desk

A repository map with abstract native-language request cards passes through a sealed testing gate

RuBench tests coding agents on native Russian repository tasks

RuBench adds 25 repository-level coding-agent tasks written natively in Russian, with withheld regression tests and product-agent runs.

The AI Feed Desk

By The AI Feed Desk

A modular agent harness surrounds a compact model core connected to tools, memory, evaluation, and runtime modules

NVIDIA says harness tuning lifts Nemotron 3 Ultra in LangChain agents

NVIDIA says tuning the LangChain Deep Agents harness for Nemotron 3 Ultra improved open-stack agent performance without retraining the model.

The AI Feed Desk

By The AI Feed Desk

A sleek abstract model engine sends fast signal trails across an engineering workbench

SpaceXAI launches Grok 4.5 for coding and agent work

SpaceXAI launched Grok 4.5 with coding benchmarks, $2 input and $6 output token pricing, and availability in Grok Build, Cursor, and its API console.

The AI Feed Desk

By The AI Feed Desk

A translucent soundwave loop splits into a live conversation path and a quieter background reasoning path

OpenAI launches GPT-Live for full-duplex ChatGPT Voice

OpenAI is rolling out GPT-Live, a full-duplex voice model family that can listen, speak, and delegate harder work in the background.

The AI Feed Desk

By The AI Feed Desk

A coding benchmark grid is inspected with several task blocks cracked or flagged

OpenAI retracts SWE-Bench Pro recommendation after benchmark audit

OpenAI audited SWE-Bench Pro and now estimates that roughly 30% of its tasks are broken, weakening a key coding-agent evaluation.

The AI Feed Desk

By The AI Feed Desk

An unfinished legal brief sits beside a rubric grid with many partial checks but only one small completed seal

Harvey LAB-AA shows legal agents still miss most complete deliverables

Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark where the top model fully passes only 14.2% of real-world legal tasks.

The AI Feed Desk

By The AI Feed Desk

Two legacy API labels move toward a dated migration checkpoint beside new V4 model cards

DeepSeek legacy API model names hit a July 24 deadline

DeepSeek says the legacy `deepseek-chat` and `deepseek-reasoner` API model names will be discontinued on July 24, pushing developers to explicit V4 model IDs.

The AI Feed Desk

By The AI Feed Desk