CapitalBench

The benchmark for AI capital allocation

Each AI model gets the same market brief, builds one portfolio in one response, and uses no browsing or tools. Real market prices determine the result.

See how AI models perform against each other, how they invest and take risk, and how they perform in the real market.

Read the CapitalBench Manifesto
Benchmark results

Which models are performing best?

Monthly and weekly tracks stay separate.

Current Monthly Benchmark

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3
GPT-5.6 Sol
Claude Opus 4.8
Grok 4.5
GPT-5.5
Gemini 3.1 Pro
Claude Opus 4.7
Claude Fable 5
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3 xAI · 4/4 scored rounds
29.5
GPT-5.6 Sol OpenAI · 4/4 scored rounds
24.2
Claude Opus 4.8 Anthropic · 4/4 scored rounds
23.4
Grok 4.5 xAI · 4/4 scored rounds
23.0
GPT-5.5 OpenAI · 4/4 scored rounds
22.9
Gemini 3.1 Pro Google · 4/4 scored rounds
20.7
Claude Opus 4.7 Anthropic · 4/4 scored rounds
20.6
Claude Fable 5 Anthropic · 4/4 scored rounds
18.9
S&P 500 S&P 500 · 4/4 scored rounds
22.0
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
4 shared resolved rounds8 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-07-15-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.3
4.15%
OpenAI GPT-5.6 Sol
3.39%
Anthropic Claude Opus 4.8
3.28%
xAI Grok 4.5
3.23%
OpenAI GPT-5.5
3.22%
Google Gemini 3.1 Pro
2.90%
Anthropic Claude Opus 4.7
2.89%
Anthropic Claude Fable 5
2.65%
S&P S&P 500
3.08%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
14.05%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.

Monthly and weekly are separate comparison tracks. Scores are never mixed across horizons. Read scoring rules

Model Risk Benchmark

Which AI models take the most risk?

Compare how aggressively each model allocates capital across its official portfolios.

Updated through August 17, 2026Higher means more risk-seeking, not better.
GPT-5.5OpenAI
79.7/100Risk-seeking
98portfolios34.1%largest position3.8%defensive assets
Grok 4.3xAI
74.3/100Risk-seeking
102portfolios44.8%largest position5.0%defensive assets
Grok 4.5xAI
71.9/100Risk-seeking
47portfolios33.3%largest position3.6%defensive assets
GPT-5.6 SolOpenAI
71.9/100Risk-seeking
43portfolios36.7%largest position8.5%defensive assets
Claude Opus 4.7Anthropic
71.3/100Risk-seeking
70portfolios31.8%largest position16.6%defensive assets
Gemini 3.1 ProGoogle
70.2/100Risk-seeking
102portfolios41.7%largest position13.1%defensive assets
Claude Opus 4.8Anthropic
69.2/100Risk-seeking
97portfolios33.9%largest position11.9%defensive assets
Claude Fable 5Anthropic
67.8/100Risk-seeking
60portfolios28.8%largest position13.8%defensive assets
Grok 4.6xAIEarly sample
65.9/100Risk-seeking
4portfolios65.0%largest position3.8%defensive assets
Claude Opus 5Anthropic
65.5/100Risk-seeking
26portfolios35.6%largest position9.4%defensive assets
Portfolio Difference

Which AI models invest most differently from the group?

Compare each model's portfolio with the choices made by the other models in the same rounds.

Updated through August 17, 2026Overall: 50% monthly / 50% weekly
Grok 4.6xAI
67.2/100
2same-round comparisonsEarly sample
Grok 4.3xAI
54.6/100
2same-round comparisonsEarly sample

Different does not mean better or worse. The score compares portfolio outputs; it does not prove copying, influence, or intent.

Explore Portfolio Difference
Performance by market

Who leads when markets rise or fall?

Monthly leaders vs. the S&P 500.

Monthly snapshot Updated Aug 14 4 of 5 market types comparable
Market fell S&P < -1.0%
Leader Grok 4.3
Average return -1.22%
Versus S&P 500 -1.90%
+0.68 pts vs S&P 6 results · Some evidence
Market rose S&P > +1.0%
Leader Grok 4.3
Average return +4.15%
Versus S&P 500 +3.08%
+1.06 pts vs S&P 4 results · Some evidence
AI positioning

What are AI models doing right now?

Live allocations before the next official score.

As of August 15, 2026
Current risk appetite As of August 15, 2026 71.5/100 Risk-seeking / Broad risk seeking
Consensus allocation As of August 15, 2026 41.3% S&P 500 (SPY) average live weight
Risk shift As of August 15, 2026 +10.5 Change vs Aug 13 portfolios
Model agreement As of August 15, 2026 Mixed 5.8 point dispersion
Current risk appetite 71.5/100 As of August 15, 2026 / Risk-seeking
As of August 15, 2026 Broad risk seeking Combined view of monthly and weekly model portfolios for this date.
Unscored portfolios 65.3/100 Separate read across every open portfolio before official scoring.
Largest current allocations
S&P 500 (SPY) 41.3% South Africa Equities (EZA) 12.5% Australia Equities (EWA) 10.6% US Small-Cap Value (IWN) 8.4% China Equities (MCHI) 6.3% Mexico Equities (EWW) 4.1%
Regime mix
Broad and cyclical equity 55.9% International equity 35.6% Defensive equity 4.1% Cash and defensive FX 2.2% Rates and credit 2.2%
Trust and proof

What makes each model test comparable?

Same brief. Same choices. One response. No tools. Real market prices.

  1. Step 1 Same report

    Every model reads the same market report.

  2. Step 2 Same choices

    Every model chooses from the same 70 assets.

  3. Step 3 One response, then locked

    No browsing, tools, or follow-up prompts. The model's submitted portfolio is frozen before results are known.

  4. Step 4 Fixed wait window

    The frozen portfolio sits untouched for 7 days or 1 month.

  5. Step 5 Prices score it

    Real ending prices decide which model did best.

Benchmark universe

What can models choose from?

The active roster, asset menu, horizons, and open rounds.

Models 10
Asset choices 70
Round lengths 2
Open rounds 21
Protocol Single-turn Non-agentic calls
Latest official results

What happened in the latest scored rounds?

Finished monthly and weekly rounds scored against real market returns.

Monthly result1 of 34
Monthly official result

Monthly result scored Aug 14

Same-window returns, ranked after final prices.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
Gemini 3.1 Pro
Grok 4.3
GPT-5.6 Sol
GPT-5.5
Grok 4.5
Claude Fable 5
Claude Opus 4.8
Claude Opus 4.7
S&P 500
XME Metals and Mining
Gemini 3.1 Pro Google
7.37%
Grok 4.3 xAI
6.64%
GPT-5.6 Sol OpenAI
6.40%
GPT-5.5 OpenAI
5.66%
Grok 4.5 xAI
5.52%
Claude Fable 5 Anthropic
5.26%
Claude Opus 4.8 Anthropic
5.23%
Claude Opus 4.7 Anthropic
3.84%
S&P 500 Benchmark
2.85%
XME Metals and Mining - Hindsight best asset
13.51%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
Gemini 3.1 Pro Google
Energy (XLE) 40% Crude Oil (USO) 30% Defense (ITA) 15% Gold (IAU) 15%
2
Grok 4.3 xAI
Energy (XLE) 50% Crude Oil (USO) 30% Financials (XLF) 20%
3
GPT-5.6 Sol OpenAI
Energy (XLE) 35% Crude Oil (USO) 25% Financials (XLF) 20% Cybersecurity (CIBR) 10% Defense (ITA) 10%
4
GPT-5.5 OpenAI
Crude Oil (USO) 45% Energy (XLE) 25% Commodities (PDBC) 15% Defense (ITA) 10% US Dollar (UUP) 5%
5
Grok 4.5 xAI
Energy (XLE) 35% Crude Oil (USO) 25% Financials (XLF) 20% Biotechnology (XBI) 10% Value (IWD) 10%
6
Claude Fable 5 Anthropic
Energy (XLE) 30% Financials (XLF) 20% Commodities (PDBC) 15% Value (IWD) 25% T-Bills (BIL) 10%
7
Claude Opus 4.8 Anthropic
Financials (XLF) 30% Energy (XLE) 20% Value (IWD) 20% Equal-Weight S&P 500 (RSP) 15% Healthcare (XLV) 15%
8
Claude Opus 4.7 Anthropic
Financials (XLF) 30% Energy (XLE) 20% Biotechnology (XBI) 15% Value (IWD) 20% T-Bills (BIL) 15%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

XME Metals and Mining - Hindsight best asset

100% Metals and Mining (XME) hindsight ceiling

Official scored round

Monthly result scored Aug 14

Audit ID: CB-2026-07-15-1M

ScoredAug 14WindowJul 15 to Aug 14Models8Asset choices70LeaderGemini 3.1 ProHorizonMonthly
Live dashboard

What is still in progress?

Open portfolios, interim returns, and upcoming score dates.

Open rounds 21 17 monthly / 4 weekly
Frozen portfolios 165 41 assets currently held
Latest close Aug 14 Live returns update before final scoring
Next score Aug 17 Official results publish after ending prices
Audit packet

How can you verify the benchmark?

Round packets expose the report, prompt, portfolios, prices, hashes, and result status.