Large Language Models News

A compact translation engine moves blank language cards between Japanese, English, and Chinese contexts

Sakana launches Namazu-powered translation in Sakana Chat

Sakana Translate adds Japanese, English, and Chinese translation to Sakana Chat, using the Namazu model series for translate, proofread, and Q&A modes.

The AI Feed Desk

By The AI Feed Desk

Generated output cards pass through a calibrated verifier gate while one unsafe stream triggers a red alarm path

Online safety monitoring paper tests calibrated verifier alarms

A new arXiv paper studies LLM safety monitoring with an external verifier signal, calibrated risk threshold, and real-time alarm decision.

The AI Feed Desk

By The AI Feed Desk

A magnifying instrument locates a highlighted region inside a transparent model-weight lattice while synthetic personal-data traces are blocked below

LACUNA asks whether LLM unlearning reaches the weights

The LACUNA testbed evaluates whether LLM unlearning methods target the parameters that stored synthetic PII, not only whether outputs stop revealing it.

The AI Feed Desk

By The AI Feed Desk

A glass model catalog panel moves from a dark retiring shelf toward a cloud model hub along a bright migration path

GitHub Models gets a July 30 shutdown date

GitHub will fully retire GitHub Models on July 30, ending the playground, model catalog, inference API, and BYOK endpoints for all customers after two July brownouts.

The AI Feed Desk

By The AI Feed Desk

A glass reasoning path highlights a trap branch while a diagnostic lens catches the error

Metacognition-Bench tests whether models notice their own mistakes

Metacognition-Bench measures whether language models detect tempting wrong reasoning paths, pairing trap-rate evaluation with adapters that flag likely free-form errors.

The AI Feed Desk

By The AI Feed Desk

A metallic chip tile connects to flowing local model tokens and a small speed gauge

BaseRT makes Apple Silicon LLM speed a runtime-design question

BaseRT's paper and release argue that native Metal kernels, unified-memory-aware layout, and custom dispatch can raise local LLM throughput on Apple Silicon.

The AI Feed Desk

By The AI Feed Desk

An open model block enters a coding model picker rail inside a cloud frame

Kimi K2.7 Code gives Copilot an open-weight model option

GitHub says Kimi K2.7 Code is the first open-weight model selectable in Copilot, with Azure hosting, gradual rollout, usage-based billing, and enterprise policy controls.

The AI Feed Desk

By The AI Feed Desk

A sequence of pull request tiles is connected by a hidden red thread under a monitoring lens

Persistent coding agents create a distributed attack surface

A new arXiv paper argues that persistent coding agents can hide malicious behavior across multiple pull requests, making monitor design a stateful problem.

The AI Feed Desk

By The AI Feed Desk

A microphone waveform passes through a fast inference core and exits as a speaker waveform

Hugging Face and Cerebras make open voice AI a latency problem

A Hugging Face and Cerebras speech-to-speech demo uses Parakeet, Gemma 4, Cerebras inference, and Qwen3TTS to show where voice AI latency actually lives.

The AI Feed Desk

By The AI Feed Desk

Abstract code blocks pass through build and deploy checkpoints before meeting a harder behavior-validation maze

ScarfBench shows coding agents still struggle with Java migrations

IBM Research's ScarfBench tests whether AI coding agents can preserve behavior while migrating Java applications across enterprise frameworks.

The AI Feed Desk

By The AI Feed Desk

A model card score links to a structured evaluation record with provenance and settings

Hugging Face and Every Eval Ever make model-card scores more inspectable

Community Evals and Every Eval Ever now connect model-page benchmark scores to structured provenance records.

The AI Feed Desk

By The AI Feed Desk

A computational biology bench with simulated genomic data paths branching into reviewed analysis decisions

OpenAI's GeneBench-Pro makes biology benchmarks about judgment

GeneBench-Pro tests whether AI agents can handle ambiguous computational-biology analysis, not just clean benchmark questions.

The AI Feed Desk

By The AI Feed Desk

A streamlined agent workspace with a smaller model core coordinating tools beside a larger reference model

Claude Sonnet 5 turns Anthropic's default model into an agent model

Anthropic made Sonnet 5 the default for Free and Pro users while positioning it as a lower-cost agentic model close to Opus 4.8.

The AI Feed Desk

By The AI Feed Desk

A coding-agent session timeline branches into tool calls, cache blocks, and latency markers

TraceLab turns real Codex and Claude Code sessions into serving data

The University of Washington TraceLab release studies coding agents as production workloads, with public traces across sessions, tool calls, tokens, cache behavior, and latency.

The AI Feed Desk

By The AI Feed Desk

A transformer-shaped lens maps scattered data points into smooth density contours

Allen AI's DiScoFormer tests one transformer for density and score

The Hugging Face writeup frames DiScoFormer as a reusable estimator for density and score, with stronger high-dimensional results than kernel density estimation.

The AI Feed Desk

By The AI Feed Desk

An AI-generated pull request passes through a software review gate with quality checkpoints

Cognition's FrontierCode asks whether AI code would survive review

FrontierCode evaluates coding agents on mergeability, code quality, scope, tests, and maintainer judgment instead of only functional correctness.

The AI Feed Desk

By The AI Feed Desk

A command-line prompt launches an inference endpoint on a small GPU cluster

Hugging Face makes vLLM serving a one-command Jobs workflow

HF Jobs can now spin up a private OpenAI-compatible vLLM endpoint for tests, evals, and batch generation without provisioning servers or managing Kubernetes.

The AI Feed Desk

By The AI Feed Desk