AI News
OpenAI launches GPT-5.6 with Sol, Terra, and Luna
OpenAI launched GPT-5.6 as three models for ChatGPT Work, Codex, and the API, adding Programmatic Tool Calling, subagents, pricing, and a new system card.
By The AI Feed Desk
RuBench tests coding agents on native Russian repository tasks
RuBench adds 25 repository-level coding-agent tasks written natively in Russian, with withheld regression tests and product-agent runs.
By The AI Feed Desk
NVIDIA says harness tuning lifts Nemotron 3 Ultra in LangChain agents
NVIDIA says tuning the LangChain Deep Agents harness for Nemotron 3 Ultra improved open-stack agent performance without retraining the model.
By The AI Feed Desk
SpaceXAI launches Grok 4.5 for coding and agent work
SpaceXAI launched Grok 4.5 with coding benchmarks, $2 input and $6 output token pricing, and availability in Grok Build, Cursor, and its API console.
By The AI Feed Desk
OpenAI launches GPT-Live for full-duplex ChatGPT Voice
OpenAI is rolling out GPT-Live, a full-duplex voice model family that can listen, speak, and delegate harder work in the background.
By The AI Feed Desk
OpenAI retracts SWE-Bench Pro recommendation after benchmark audit
OpenAI audited SWE-Bench Pro and now estimates that roughly 30% of its tasks are broken, weakening a key coding-agent evaluation.
By The AI Feed Desk
Hugging Face cloud handoffs turn model pages into deployment paths
AWS and SkyPilot integrations show Hugging Face model pages becoming handoffs into managed SageMaker workflows and portable multi-cloud GPU jobs.
By The AI Feed Desk
Hugging Face models reach Microsoft Foundry Managed Compute
Hugging Face and Microsoft are putting curated open-weight models on Foundry Managed Compute with weekly refreshes, Azure-staged weights, and scanned runtimes.
By The AI Feed Desk
GitHub brings Kimi K2.7 to Copilot Business and Enterprise
GitHub added Kimi K2.7 Code to Copilot Business and Enterprise while expanding per-user AI budgets and review-cycle metrics.
By The AI Feed Desk
GitHub Copilot app reaches every plan
GitHub made the Copilot desktop app available across Copilot Free, GitHub Education, paid plans, and BYOK sessions without a Copilot subscription.
By The AI Feed Desk
Harvey LAB-AA shows legal agents still miss most complete deliverables
Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark where the top model fully passes only 14.2% of real-world legal tasks.
By The AI Feed Desk
Hugging Face LeRobot 0.6 closes the robot learning loop
LeRobot v0.6.0 adds world-model policies, reward models, simulation benchmarks, rollout tooling, depth data, and HF Jobs training to Hugging Face's robotics stack.
By The AI Feed Desk
Alberta used Claude Code to scan 466 million lines of government code
Anthropic says Alberta used Claude Code agents to review legacy government systems, find vulnerabilities, generate fixes, and build continuous security-review agents.
By The AI Feed Desk
Databricks frames OpenAI partnership around production agents
Databricks' DAIS 2026 recap presents OpenAI as the intelligence layer for enterprise agents while Databricks owns context, governance, and production data workflows.
By The AI Feed Desk
DeepSeek legacy API model names hit a July 24 deadline
DeepSeek says the legacy `deepseek-chat` and `deepseek-reasoner` API model names will be discontinued on July 24, pushing developers to explicit V4 model IDs.
By The AI Feed Desk
AutomationBench-AA shows agents still break business guardrails
Artificial Analysis and Zapier launched AutomationBench-AA, a SaaS-agent benchmark that scores completed objectives only when business guardrails stay intact.
By The AI Feed Desk
Program-as-Weights turns fuzzy functions into local adapters
A new arXiv paper proposes compiling natural-language fuzzy functions into small local adapters, shifting some LLM work from repeated API calls to reusable artifacts.
By The AI Feed Desk



















