A repository map with abstract native-language request cards passes through a sealed testing gate
A repository map with abstract native-language request cards passes through a sealed testing gate
+ Large Language Models News

RuBench tests coding agents on native Russian repository tasks

RuBench adds 25 repository-level coding-agent tasks written natively in Russian, with withheld regression tests and product-agent runs.

RuBench is a new repository-level coding-agent benchmark built around tasks written natively in Russian rather than translated from English.

The arXiv paper, submitted July 7, describes RuBench 1.0 as 25 tasks mined from recent fix commits in five live open-source repositories: aiohttp, aiogram, Laravel, NestJS, and Fastify. The tasks span Python, PHP, TypeScript, and JavaScript.

The paper says each task statement was written from scratch in Russian in the style of a customer request. The benchmark withholds the upstream maintainer regression tests used for grading and publishes a SHA-256 manifest for the grading oracles.

The benchmark is small, but the setup is useful

Most repository-level coding benchmarks are English by design. That leaves a gap for teams whose users, support tickets, or product requirements arrive in another language.

RuBench tries to measure that setting directly. It evaluates product configurations, not just raw model APIs: Claude Code with Opus 4.8, Sonnet 5, and Haiku 4.5, plus Codex CLI with GPT-5.5, with three independent runs each.

The paper reports that the best configuration resolves 78.7% of tasks. It is careful about the sample size: at 25 tasks, only gaps to the weakest model are statistically resolvable.

That caveat is important. RuBench should not be treated as a general model ranking. It is better read as a design example for multilingual, product-level coding-agent tests.

Product behavior can change what is measured

The paper’s most practical finding is about a fifth, hors-concours configuration: Claude Code with Fable 5.

The author says trajectory audits found that on 5 of 25 tasks, or 20%, an official safeguard fallback silently rerouted routine HTTP-protocol fixes to Opus 4.8. That means the deployed product, not only the named model, became the measured unit.

That distinction matters for any coding-agent benchmark. A user may think they are evaluating a model, while the product is also applying routing, safeguards, hidden policies, or fallback behavior.

For teams buying coding agents, that means model labels are not enough. They need run logs, routing transparency, and repeatable evaluation conditions.

Sources

The AI Feed Desk

The AI Feed Desk

Editorial desk

The AI Feed Desk tracks AI provider updates, model releases, agent tooling, and enterprise adoption, turning fast-moving announcements into source-linked context for builders and operators.

Noticed a typo, incorrect information, or translation error?

Tell us so we can fix it.

Help Improve This Article

Related Articles

An AI-generated pull request passes through a software review gate with quality checkpoints

Cognition's FrontierCode asks whether AI code would survive review

FrontierCode evaluates coding agents on mergeability, code quality, scope, tests, and maintainer judgment instead of only functional correctness.

The AI Feed Desk

By The AI Feed Desk

A coding benchmark grid is inspected with several task blocks cracked or flagged

OpenAI retracts SWE-Bench Pro recommendation after benchmark audit

OpenAI audited SWE-Bench Pro and now estimates that roughly 30% of its tasks are broken, weakening a key coding-agent evaluation.

The AI Feed Desk

By The AI Feed Desk

A business briefcase opens into documents, messages, and evaluation scorecards

AA-Briefcase tests agents on messy business work

Artificial Analysis' AA-Briefcase benchmark evaluates models on multi-week knowledge-work projects with documents, email, Slack data, deliverables, and graded analysis quality.

The AI Feed Desk

By The AI Feed Desk

An unfinished legal brief sits beside a rubric grid with many partial checks but only one small completed seal

Harvey LAB-AA shows legal agents still miss most complete deliverables

Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark where the top model fully passes only 14.2% of real-world legal tasks.

The AI Feed Desk

By The AI Feed Desk

A coding benchmark workbench with agent paths, repository blocks, cost meters, and verification checkmarks

DeepSWE makes coding-agent rankings a cost question

DeepSWE's June 20 leaderboard update separates frontier coding agents by pass rate, cost, output tokens, and agent steps across long-horizon software tasks.

The AI Feed Desk

By The AI Feed Desk