Arito Ranks #1 on Spreadsheet Bench v2

Arito ranks #1 on SpreadsheetBench v2 with a 45.46% verified score. Learn how Arito’s purpose-built AI agent harness, Excel tooling, visual feedback, and context engineering help frontier models perform complex spreadsheet tasks.

How a purpose-built agent harness unlocks more from frontier models.

Arito is now #1 on SpreadsheetBench v2 and the most interesting part isn’t the model.

As frontier models improve, it’s tempting to think the agent harness matters less: plug the latest model into a generic scaffold, and most of the gains come along for free.

Our verified SpreadsheetBench v2 result suggests otherwise.

SpreadsheetBench v2 evaluates end-to-end business spreadsheet workflows: financial modeling, template filling, KPI generation, charts, and more.

πŸ”Ή Arito: 45.46% – verified, #1
πŸ”Ή Fable 5: 34.70%
πŸ”Ή GPT-5.6 Sol: 32.40%
πŸ”Ή Opus 4.8: 31.60%

That’s a 30%+ relative improvement over frontier models alone. And here’s the thing: we run on a frontier model too. The entire delta is the harness.

The difference comes from the system around the model.
What moved the needle:

Purpose-built Excel tooling
A dedicated toolset for parsing, editing, and reasoning over spreadsheets, so the agent isn’t reinventing workbook manipulation on every task.

Visual feedback
The agent takes screenshots before editing to understand the workbook’s structure, and again afterward to verify the result. Spreadsheets are a visual medium; an agent that can’t see them is working blindfolded.

The real Excel calculation engine in the loop
Via Microsoft Graph API, the agent can inspect actual calculated results rather than simulate them. It can write a formula, see what Excel really computes, fix mistakes, and revert to previous revisions when needed.

Advanced context engineering
Automatic context management, compaction, and working memory. Deep spreadsheet tasks are long-horizon and highly stateful: hundreds of interdependent edits across dozens of sheets. Without disciplined context management, agents lose track of the work long before they finish it.

None of this replaces a strong model.

But for deep, domain-specific work: a 40-sheet financial model, a workbook with hundreds of broken formulas, the model alone isn’t enough.

The model sets the ceiling. The harness determines how much of that capability you can actually use.

TRY ARITO NOW

Ready for continuous identity protection without gaps or guesswork?

β†’

About The Author

Michael Estrin

Michael Estrin is CTO and co-founder of arito. He is a second-time founder and technology leader with deep expertise in scalable systems, artificial intelligence, machine learning, and infrastructure. He previously co-founded LEVL, where he served as CTO from inception through its acquisition by Comcast and Charter in 2022, leading the company’s technology strategy and building its core platform. Earlier in his career, Michael served in Israeli intelligence and held engineering roles at Google and Dell. Today, he focuses on building the core infrastructure behind arito’s enterprise AI platform, designing reliable, scalable systems that enable AI agents to operate on complex business data and workflows, starting with finance.