Independent research on the systems everyone is building

Montana Research Foundation is a home for scientists and researchers working on the science of AI, and sharing what they find.

Get in touch

The foundation

Research is what we do

We spend our days on AI and ML research: asking good questions, running careful experiments, and publishing what we learn so anyone can build on it.

Independent

We're a non-profit, which means we get to pick research questions on their merit and follow them wherever they lead.

Rigorous

Every piece of work passes an internal editorial and reproducibility review before it goes out.

Open

We publish what we learn so the wider research community can build on it.

We chase the questions
too interesting to leave unanswered.

And everything we find gets published openly, for anyone who wants to use it.

Reports

Findings, written to be acted on

All reports
MRF-R-2026-03 Report · August 2026

Agents in Financial Services

Three quarters of UK financial firms already use AI, a third say they fully understand the systems they run, and more than a third of the US population has dealt with a bank's chatbot. Against that, a tribunal has held an airline liable for its chatbot's advice, a finance question-answering benchmark found a frontier model wrong or silent on four questions in five, and the EU has classified credit scoring as high-risk. This report puts the numbers together and sets out controls a risk function can sign off.

Read the report
MRF-R-2026-02 Report · August 2026

Coding Agents in Production

Engineering leaders are deciding how far to lean on coding agents on the strength of vendor demos and leaderboard scores. The published evidence is more mixed than either suggests: two controlled trials point in opposite directions, delivery telemetry shows stability falling as adoption rises, and the benchmark most often quoted has been retired by one of its own maintainers. This report sets out that evidence and the measurements a team should run on its own codebase before it changes how it works.

Read the report
MRF-R-2026-01 Report · August 2026

The State of Agent Evaluation

Three foundation preprints asked how frontier agents are measured once the grading criterion is held out of the environment. This report draws their conclusions together for lab and policy readers: which classes of broken grader a reference solution cannot catch, how far declared confidence sits above measured generalisation on certified task families, and how one sentence of ability framing moved the reasoning effort of two frontier configurations. Every figure traces to a committed run record.

Read the report

This might be your kind of place.

Get in touch