Artificial Analysis launched Harvey LAB-AA on July 7, a legal-agent benchmark built from 120 private tasks created by Harvey across 24 legal practice areas.
The result is a useful brake on legal-agent hype. The top system, Claude Fable 5 with an Opus 4.8 fallback, fully passes 14.2% of tasks. Claude Opus 4.8 and GLM-5.2 tie next at 7.5%. Artificial Analysis says 13 of 28 evaluated models fully pass zero tasks.
The benchmark’s primary score is strict. A task counts only when every binary criterion in the rubric passes. That matters for legal work because a deliverable with nine correct parts and one missed requirement can still be unusable.
The gap is between partial correctness and finished work
Artificial Analysis also reports criterion pass rates. On that looser view, four models clear 90%: Claude Fable 5 at 93.6%, Claude Opus 4.8 at 91.1%, GLM-5.2 at 91.0%, and Claude Sonnet 5 at 90.1%.
That spread is the story. Legal agents can hit many local requirements inside a task while still missing enough of the complete deliverable to fail the professional standard.
Cost complicates the ranking. Artificial Analysis says Claude Fable 5 costs about $18.90 per task in the benchmark. Claude Sonnet 5 costs about $11.80, Claude Opus 4.8 about $8.20, and GLM-5.2 about $1.30 while tying Opus 4.8 on all-pass rate.
That does not make GLM-5.2 the universal legal choice. It does mean buyers should compare full-task success against cost, latency, review burden, and the legal domain being tested.
The benchmark should change how teams pilot legal AI
The practical lesson is to stop measuring only whether a model sounds legally fluent. Teams need pass/fail rubrics tied to actual output requirements: did the agent identify the right documents, apply the right jurisdictional frame, cite the right clause, flag the right caveat, and produce the requested form.
Harvey LAB-AA is private, so outside teams cannot reproduce every task directly. But its scoring design is portable. A legal team can create smaller internal task sets with binary criteria and require full deliverable pass rates before expanding deployment.
The counter-case is that private benchmarks can overfit to one vendor’s task design and rubric choices. The result should be read as a benchmark signal, not a complete map of legal-agent quality.





