A business briefcase opens into documents, messages, and evaluation scorecards
A business briefcase opens into documents, messages, and evaluation scorecards
+ Large Language Models News

AA-Briefcase tests agents on messy business work

Artificial Analysis' AA-Briefcase benchmark evaluates models on multi-week knowledge-work projects with documents, email, Slack data, deliverables, and graded analysis quality.

Artificial Analysis has published AA-Briefcase, an agentic knowledge-work benchmark built around realistic business projects rather than single prompts. The benchmark evaluates models across four multi-week projects and 91 tasks, using nearly 2,000 source files, more than 3,500 emails, and 25,000 Slack messages.

That structure is the story. Agent benchmarks are moving toward the shape of actual office work: messy context, conflicting inputs, multiple deliverables, and judgments about whether the answer is correct, analytical, and presentable.

AA-Briefcase is not just asking whether a model can produce a polished response. It asks whether the model can navigate a project folder.

The benchmark is closer to how agents are sold

Most enterprise agent promises are not about trivia answers. They are about knowledge work: read the documents, find the relevant thread, understand the spreadsheet, produce the memo, prepare the slide, reconcile the contradictions, and finish the task.

AA-Briefcase tries to simulate that environment. Artificial Analysis says the scenarios are multi-week workflows in data science, product management, and corporate strategy. The tasks were built by experts from companies including Google, McKinsey & Company, and Boston Consulting Group.

That matters because the unit of work is a deliverable, not a chat response. A model can sound competent in a short answer while still missing the hidden constraint in an email, misunderstanding a Slack discussion, or producing a presentation that looks good but fails the rubric.

The grading mixes correctness and quality

Artificial Analysis describes AA-Briefcase as combining rubric checks with pairwise grading. The rubric side tests verifiable task success. The pairwise side evaluates analytical quality and presentation quality.

That combination is useful because business work has more than one failure mode. A model can be factually wrong. It can be analytically weak. It can present the right answer in a way that a stakeholder cannot use. It can produce something polished but unsupported.

The benchmark’s combined AA-Briefcase Elo aggregates rubric pass rate, analytical quality Elo, and presentation Elo. That makes the score more holistic than a pure pass/fail metric. It also means readers should be careful: an Elo score is a benchmark-specific measurement, not a universal promise that one model will perform best inside a particular company.

Long-running tasks create a cost question

Artificial Analysis followed the benchmark launch with a time-per-task article. That is the right next question. Long-horizon agents do not only differ in quality; they differ in how long they run, how many steps they take, and how much they cost per deliverable.

For enterprise buyers, that is often the real comparison. A model that produces better analysis but takes much longer may still be worth it for a strategy memo. The same trade-off may be unacceptable for routine operations work. A cheaper model that is “good enough” for structured tasks may beat a frontier model on cost-performance.

That is why the benchmark’s business-work framing matters. It gives teams a way to discuss agent performance in units closer to work output: tasks, deliverables, time, and quality.

Sources

The AI Feed Desk

The AI Feed Desk

Editorial desk

The AI Feed Desk tracks AI provider updates, model releases, agent tooling, and enterprise adoption, turning fast-moving announcements into source-linked context for builders and operators.

Noticed a typo, incorrect information, or translation error?

Tell us so we can fix it.

Help Improve This Article

Related Articles

An unfinished legal brief sits beside a rubric grid with many partial checks but only one small completed seal

Harvey LAB-AA shows legal agents still miss most complete deliverables

Artificial Analysis launched Harvey LAB-AA, a legal-agent benchmark where the top model fully passes only 14.2% of real-world legal tasks.

The AI Feed Desk

By The AI Feed Desk

An AI agent moves tasks across business apps while warning rails light up around the workflow

AutomationBench-AA shows agents still break business guardrails

Artificial Analysis and Zapier launched AutomationBench-AA, a SaaS-agent benchmark that scores completed objectives only when business guardrails stay intact.

The AI Feed Desk

By The AI Feed Desk

An AI-generated pull request passes through a software review gate with quality checkpoints

Cognition's FrontierCode asks whether AI code would survive review

FrontierCode evaluates coding agents on mergeability, code quality, scope, tests, and maintainer judgment instead of only functional correctness.

The AI Feed Desk

By The AI Feed Desk

A testing gauge compares a clean tool path with a longer tangled debugging path

Hugging Face measures whether tools are agent-friendly

Hugging Face's agent-focused benchmark tests whether software changes help coding agents finish tasks with fewer errors, tokens, and detours.

The AI Feed Desk

By The AI Feed Desk

A model card score links to a structured evaluation record with provenance and settings

Hugging Face and Every Eval Ever make model-card scores more inspectable

Community Evals and Every Eval Ever now connect model-page benchmark scores to structured provenance records.

The AI Feed Desk

By The AI Feed Desk