FrontierFinance is a finance AI agent benchmark testing systems on 220 real investment queries, built entirely by Samaya AI's finance experts.
Photo source:
samaya.ai
Most
tests built to measure how good an AI system is at finance only check one
thing: can it pull a number out of a filing. Samaya AI, a company built by
former Google Brain researcher Maithra Raghu, argues that misses most of what
an investor actually does all day. Its answer is FrontierFinance, a finance AI
agent benchmark built entirely from real investment workflows rather than
simple data-extraction tasks, covering everything from early-stage idea
screening to ongoing portfolio monitoring.
Samaya's
own agent platform, the product this benchmark was built to test and improve,
is already used by named clients including Morgan Stanley, where Global
Director of Research Katy Huberty has credited the system with powering
research analysis across the firm's Institutional Securities Group. That
existing deployment is part of what makes the benchmark worth taking seriously:
it wasn't built in the abstract, but by a team already running this software
inside real investment firms.
FrontierFinance
contains 220 expert-crafted queries paired with 11,543 grading rubrics,
covering six distinct finance use cases rather than one narrow task type. Every
query and its rubric were built through a four-stage process led entirely by
in-house finance experts: drafting realistic, sometimes ambiguous questions
grounded in everyday investor work, writing the specific rubric points a good
answer needs to satisfy, reviewing and removing anything redundant or
subjective, and finally rebalancing the set across use cases and difficulty
levels. That process started from a much larger pool of more than 4,500 fully
annotated queries, with the final 220 selected specifically to form a balanced
public benchmark.
Every
rubric point traces back to a real, publicly available source rather than an
invented answer key. SEC filings supply the largest single share at 39 percent,
with the remainder pulled from company transcripts and investor presentations,
general professional finance knowledge, and market data. According to Samaya,
this sourcing mix shifts depending on the use case being tested, reflecting how
differently an investor might approach, say, screening a new company versus
tracking a portfolio holding they already own.
To
back up the claim that FrontierFinance is genuinely harder than existing
finance benchmarks, Samaya used a comparison method borrowed from competitive
game rating systems. Every example from FrontierFinance and from comparison
benchmarks was pooled together, and an AI judge compared pairs of examples
head-to-head to determine which was harder to fully answer. Those pairwise
comparisons were then converted into a difficulty score for each query, the
same statistical approach used in systems like chess Elo ratings. The result,
according to Samaya, shows FrontierFinance covering a wider range of difficulty
with a higher average difficulty than existing public finance benchmarks.
Testing
itself runs across three distinct setups, called harnesses: a plain AI model
paired only with basic web search, an open-source setup called Finance Agent v2
that adds six specialized tools including the SEC's EDGAR filing database and a
dedicated market data API, and Samaya's own in-house harness combining its
custom models with its own data retrieval systems. Every setup is scored on the
same three measures: how many rubric points an answer satisfies, judged by
majority vote across three separate AI judges; the average cost of a single
query; and how long a query takes to answer.
Samaya's
own in-house agent system scored 50.8 percent on FrontierFinance's rubric
satisfaction measure, the highest of any system tested, ahead of Claude Fable 5
at 49.2 percent, Claude Opus 4.8 at 45.0 percent, and GPT-5.5 at 43.5 percent.
According to Samaya, its system achieved that top score at roughly a quarter of
the typical cost of running Claude Fable 5 on the same tasks. Among freely
available open-weight models, DeepSeek V4 Pro reached 40.5 percent, close
behind the more expensive commercial systems while costing a fraction as much
to run.
Breaking
results down by category surfaced a consistent pattern: every system tested,
including Samaya's own, struggled most with screening and discovery tasks and
with broad sector, industry, and macroeconomic questions, while performing
comparatively better on tasks focused on formatting and presenting an answer
clearly. That gap between formatting quality and genuine analytical depth is
arguably the more interesting finding than any single leaderboard score, since
it points to where current AI systems are still weakest at the actual
judgment-heavy parts of investment work.
FrontierFinance
is a genuinely more comprehensive test than benchmarks focused narrowly on
pulling numbers from filings, and Samaya has made the underlying dataset and
grading code publicly available for anyone to check independently. That
openness matters, because Samaya is also the company whose own commercial
product came out on top of the results it published, a real point worth being
direct about rather than glossing over. A benchmark's creator scoring best on
its own test isn't evidence of bias by itself, but it does mean outside
verification, now possible thanks to the public release, is the part that
actually settles the question rather than the announcement alone.
Even
setting that aside, the headline result is a reminder of how far finance AI
still has to go: the best-performing system on FrontierFinance satisfied barely
half of the expert-written rubric points across the benchmark's 220 queries.
Whatever system comes out on top, a technology only just crossing the halfway
mark on realistic investment questions is still early in proving it can be
trusted with the highest-stakes, judgment-heavy parts of the job.
Please subscribe to have unlimited access to our innovations.