An unfinished legal brief sits beside a rubric grid with many partial checks but only one small completed seal
An unfinished legal brief sits beside a rubric grid with many partial checks but only one small completed seal
+ Large Language Models News

Harvey LAB-AA shows legal agents still miss most complete deliverables

Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark where the top model fully passes only 14.2% of real-world legal tasks.

Artificial Analysis launched Harvey LAB-AA on July 7, a legal-agent benchmark built from 120 private tasks created by Harvey across 24 legal practice areas.

The result is a useful brake on legal-agent hype. The top system, Claude Fable 5 with an Opus 4.8 fallback, fully passes 14.2% of tasks. Claude Opus 4.8 and GLM-5.2 tie next at 7.5%. Artificial Analysis says 13 of 28 evaluated models fully pass zero tasks.

The benchmark’s primary score is strict. A task counts only when every binary criterion in the rubric passes. That matters for legal work because a deliverable with nine correct parts and one missed requirement can still be unusable.

The gap is between partial correctness and finished work

Artificial Analysis also reports criterion pass rates. On that looser view, four models clear 90%: Claude Fable 5 at 93.6%, Claude Opus 4.8 at 91.1%, GLM-5.2 at 91.0%, and Claude Sonnet 5 at 90.1%.

That spread is the story. Legal agents can hit many local requirements inside a task while still missing enough of the complete deliverable to fail the professional standard.

Cost complicates the ranking. Artificial Analysis says Claude Fable 5 costs about $18.90 per task in the benchmark. Claude Sonnet 5 costs about $11.80, Claude Opus 4.8 about $8.20, and GLM-5.2 about $1.30 while tying Opus 4.8 on all-pass rate.

That does not make GLM-5.2 the universal legal choice. It does mean buyers should compare full-task success against cost, latency, review burden, and the legal domain being tested.

The practical lesson is to stop measuring only whether a model sounds legally fluent. Teams need pass/fail rubrics tied to actual output requirements: did the agent identify the right documents, apply the right jurisdictional frame, cite the right clause, flag the right caveat, and produce the requested form.

Harvey LAB-AA is private, so outside teams cannot reproduce every task directly. But its scoring design is portable. A legal team can create smaller internal task sets with binary criteria and require full deliverable pass rates before expanding deployment.

The counter-case is that private benchmarks can overfit to one vendor’s task design and rubric choices. The result should be read as a benchmark signal, not a complete map of legal-agent quality.

Sources

The AI Feed Desk

The AI Feed Desk

Editorial desk

The AI Feed Desk tracks AI provider updates, model releases, agent tooling, and enterprise adoption, turning fast-moving announcements into source-linked context for builders and operators.

Noticed a typo, incorrect information, or translation error?

Tell us so we can fix it.

Help Improve This Article

Related Articles

A business briefcase opens into documents, messages, and evaluation scorecards

AA-Briefcase tests agents on messy business work

Artificial Analysis' AA-Briefcase benchmark evaluates models on multi-week knowledge-work projects with documents, email, Slack data, deliverables, and graded analysis quality.

The AI Feed Desk

By The AI Feed Desk

An AI agent moves tasks across business apps while warning rails light up around the workflow

AutomationBench-AA shows agents still break business guardrails

Artificial Analysis and Zapier launched AutomationBench-AA, a SaaS-agent benchmark that scores completed objectives only when business guardrails stay intact.

The AI Feed Desk

By The AI Feed Desk

An AI-generated pull request passes through a software review gate with quality checkpoints

Cognition's FrontierCode asks whether AI code would survive review

FrontierCode evaluates coding agents on mergeability, code quality, scope, tests, and maintainer judgment instead of only functional correctness.

The AI Feed Desk

By The AI Feed Desk

A testing gauge compares a clean tool path with a longer tangled debugging path

Hugging Face measures whether tools are agent-friendly

Hugging Face's agent-focused benchmark tests whether software changes help coding agents finish tasks with fewer errors, tokens, and detours.

The AI Feed Desk

By The AI Feed Desk

A model card score links to a structured evaluation record with provenance and settings

Hugging Face and Every Eval Ever make model-card scores more inspectable

Community Evals and Every Eval Ever now connect model-page benchmark scores to structured provenance records.

The AI Feed Desk

By The AI Feed Desk