FrontierFinance Is a Finance AI Agent Benchmark for Real Investors

FrontierFinance is a finance AI agent benchmark testing systems on 220 real investment queries, built entirely by Samaya AI's finance experts.

Photo source:

samaya.ai

FrontierFinance Is a Finance AI Agent Benchmark for Real Investors

Most tests built to measure how good an AI system is at finance only check one thing: can it pull a number out of a filing. Samaya AI, a company built by former Google Brain researcher Maithra Raghu, argues that misses most of what an investor actually does all day. Its answer is FrontierFinance, a finance AI agent benchmark built entirely from real investment workflows rather than simple data-extraction tasks, covering everything from early-stage idea screening to ongoing portfolio monitoring.

Samaya's own agent platform, the product this benchmark was built to test and improve, is already used by named clients including Morgan Stanley, where Global Director of Research Katy Huberty has credited the system with powering research analysis across the firm's Institutional Securities Group. That existing deployment is part of what makes the benchmark worth taking seriously: it wasn't built in the abstract, but by a team already running this software inside real investment firms.

Built From Real Investor Questions, Not Just Data Pulls

FrontierFinance contains 220 expert-crafted queries paired with 11,543 grading rubrics, covering six distinct finance use cases rather than one narrow task type. Every query and its rubric were built through a four-stage process led entirely by in-house finance experts: drafting realistic, sometimes ambiguous questions grounded in everyday investor work, writing the specific rubric points a good answer needs to satisfy, reviewing and removing anything redundant or subjective, and finally rebalancing the set across use cases and difficulty levels. That process started from a much larger pool of more than 4,500 fully annotated queries, with the final 220 selected specifically to form a balanced public benchmark.

Every rubric point traces back to a real, publicly available source rather than an invented answer key. SEC filings supply the largest single share at 39 percent, with the remainder pulled from company transcripts and investor presentations, general professional finance knowledge, and market data. According to Samaya, this sourcing mix shifts depending on the use case being tested, reflecting how differently an investor might approach, say, screening a new company versus tracking a portfolio holding they already own.

Testing Difficulty the Same Way Chess Ratings Work

To back up the claim that FrontierFinance is genuinely harder than existing finance benchmarks, Samaya used a comparison method borrowed from competitive game rating systems. Every example from FrontierFinance and from comparison benchmarks was pooled together, and an AI judge compared pairs of examples head-to-head to determine which was harder to fully answer. Those pairwise comparisons were then converted into a difficulty score for each query, the same statistical approach used in systems like chess Elo ratings. The result, according to Samaya, shows FrontierFinance covering a wider range of difficulty with a higher average difficulty than existing public finance benchmarks.

Testing itself runs across three distinct setups, called harnesses: a plain AI model paired only with basic web search, an open-source setup called Finance Agent v2 that adds six specialized tools including the SEC's EDGAR filing database and a dedicated market data API, and Samaya's own in-house harness combining its custom models with its own data retrieval systems. Every setup is scored on the same three measures: how many rubric points an answer satisfies, judged by majority vote across three separate AI judges; the average cost of a single query; and how long a query takes to answer.

Where the Numbers Actually Landed

Samaya's own in-house agent system scored 50.8 percent on FrontierFinance's rubric satisfaction measure, the highest of any system tested, ahead of Claude Fable 5 at 49.2 percent, Claude Opus 4.8 at 45.0 percent, and GPT-5.5 at 43.5 percent. According to Samaya, its system achieved that top score at roughly a quarter of the typical cost of running Claude Fable 5 on the same tasks. Among freely available open-weight models, DeepSeek V4 Pro reached 40.5 percent, close behind the more expensive commercial systems while costing a fraction as much to run.

Breaking results down by category surfaced a consistent pattern: every system tested, including Samaya's own, struggled most with screening and discovery tasks and with broad sector, industry, and macroeconomic questions, while performing comparatively better on tasks focused on formatting and presenting an answer clearly. That gap between formatting quality and genuine analytical depth is arguably the more interesting finding than any single leaderboard score, since it points to where current AI systems are still weakest at the actual judgment-heavy parts of investment work.

A Benchmark Built and Graded by the Same Company

FrontierFinance is a genuinely more comprehensive test than benchmarks focused narrowly on pulling numbers from filings, and Samaya has made the underlying dataset and grading code publicly available for anyone to check independently. That openness matters, because Samaya is also the company whose own commercial product came out on top of the results it published, a real point worth being direct about rather than glossing over. A benchmark's creator scoring best on its own test isn't evidence of bias by itself, but it does mean outside verification, now possible thanks to the public release, is the part that actually settles the question rather than the announcement alone.

Even setting that aside, the headline result is a reminder of how far finance AI still has to go: the best-performing system on FrontierFinance satisfied barely half of the expert-written rubric points across the benchmark's 220 queries. Whatever system comes out on top, a technology only just crossing the halfway mark on realistic investment questions is still early in proving it can be trusted with the highest-stakes, judgment-heavy parts of the job.

Lock

You have exceeded your free limits for viewing our premium content

Please subscribe to have unlimited access to our innovations.