CapitalBench

The benchmark for AI capital allocation

Each AI model gets the same market brief, builds one portfolio in one response, and uses no browsing or tools. Real market prices determine the result.

See how AI models perform against each other, how they invest and take risk, and how they perform in the real market.

Read the CapitalBench Manifesto
Benchmark results

Which models are performing best?

Monthly and weekly tracks stay separate.

Current Monthly Benchmark

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Claude Fable 5
GPT-5.6 Sol
Claude Opus 5
Grok 4.3
Grok 4.6
Grok 4.5
Gemini 3.1 Pro
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Claude Fable 5 Anthropic · 12/12 scored rounds
3.8
GPT-5.6 Sol OpenAI · 12/12 scored rounds
-0.9
Claude Opus 5 Anthropic · 12/12 scored rounds
-4.7
Grok 4.3 xAI · 12/12 scored rounds
-5.6
Grok 4.6 xAI · 12/12 scored rounds
-6.4
Grok 4.5 xAI · 12/12 scored rounds
-6.8
Gemini 3.1 Pro Google · 12/12 scored rounds
-13.4
S&P 500 S&P 500 · 12/12 scored rounds
-1.1
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
12 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-30-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

Anthropic Claude Fable 5
0.79%
OpenAI GPT-5.6 Sol
-0.19%
Anthropic Claude Opus 5
-0.98%
xAI Grok 4.3
-1.16%
xAI Grok 4.6
-1.33%
xAI Grok 4.5
-1.42%
Google Gemini 3.1 Pro
-2.79%
S&P S&P 500
-0.22%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
20.83%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.

Monthly and weekly are separate comparison tracks. Scores are never mixed across horizons. Read scoring rules

Model Risk Benchmark

Which AI models take the most risk?

Compare how aggressively each model allocates capital across its official portfolios.

Updated through September 30, 2026Higher means more risk-seeking, not better.
GPT-5.5OpenAI
79.7/100Risk-seeking
98portfolios34.1%largest position3.8%defensive assets
Claude Fable 5.1Anthropic
74.0/100Risk-seeking
22portfolios46.4%largest position6.1%defensive assets
GPT-5.6 SolOpenAI
73.6/100Risk-seeking
69portfolios36.1%largest position8.1%defensive assets
Grok 4.5xAI
73.3/100Risk-seeking
93portfolios34.8%largest position6.8%defensive assets
Grok 4.3xAI
72.6/100Risk-seeking
148portfolios49.7%largest position6.1%defensive assets
Claude Opus 5Anthropic
72.3/100Risk-seeking
68portfolios35.7%largest position7.5%defensive assets
Gemini 3.1 ProGoogle
71.4/100Risk-seeking
148portfolios39.8%largest position13.6%defensive assets
Claude Opus 4.7Anthropic
71.3/100Risk-seeking
70portfolios31.8%largest position16.6%defensive assets
Claude Fable 5Anthropic
70.8/100Risk-seeking
84portfolios31.3%largest position12.1%defensive assets
Grok 4.6xAI
70.7/100Risk-seeking
50portfolios65.3%largest position5.1%defensive assets
GPT-6 AstraOpenAI
69.6/100Risk-seeking
20portfolios42.8%largest position8.8%defensive assets
Claude Opus 4.8Anthropic
69.2/100Risk-seeking
99portfolios35.3%largest position11.7%defensive assets
Portfolio Difference

Which AI models invest most differently from the group?

Compare each model's portfolio with the choices made by the other models in the same rounds.

Updated through September 30, 2026Overall: 50% monthly / 50% weekly
Grok 4.6xAI
59.2/100
48same-round comparisonsEstablished sample
Grok 4.5xAI
53.5/100
48same-round comparisonsEstablished sample
GPT-6 AstraOpenAI
51.0/100
20same-round comparisonsEstablished sample

Different does not mean better or worse. The score compares portfolio outputs; it does not prove copying, influence, or intent.

Explore Portfolio Difference
Performance by market

Who leads when markets rise or fall?

Monthly leaders vs. the S&P 500.

Monthly snapshot Updated Sep 30 4 of 5 market types comparable
Market fell S&P < -1.0%
Leader Grok 4.3
Average return -1.22%
Versus S&P 500 -1.90%
+0.68 pts vs S&P 6 results · Some evidence
Market rose S&P > +1.0%
Leader Grok 4.5
Average return +4.11%
Versus S&P 500 +3.90%
+0.22 pts vs S&P 6 results · Some evidence
AI positioning

What are AI models doing right now?

Live allocations before the next official score.

As of October 1, 2026
Current risk appetite As of October 1, 2026 68.0/100 Risk-seeking / Broad risk seeking
Consensus allocation As of October 1, 2026 56.4% S&P 500 (SPY) average live weight
Risk shift As of October 1, 2026 -2.8 Change vs Sep 28 portfolios
Model agreement As of October 1, 2026 Tight 2.5 point dispersion
Current risk appetite 68.0/100 As of October 1, 2026 / Risk-seeking
As of October 1, 2026 Broad risk seeking Combined view of monthly and weekly model portfolios for this date.
Unscored portfolios 72.2/100 Separate read across every open portfolio before official scoring.
Largest current allocations
S&P 500 (SPY) 56.4% Silver (SLV) 7.5% Crude Oil (USO) 5.0% Financials Sector (XLF) 5.0% Metals and Mining (XME) 5.0% US Large-Cap Value (IWD) 5.0%
Regime mix
Broad and cyclical equity 68.6% Real assets and inflation 20.0% Growth and technology 4.6% International equity 4.3% Rates and credit 2.5%
Trust and proof

What makes each model test comparable?

Same brief. Same choices. One response. No tools. Real market prices.

  1. Step 1 Same report

    Every model reads the same market report.

  2. Step 2 Same choices

    Every model chooses from the same 70 assets.

  3. Step 3 One response, then locked

    No browsing, tools, or follow-up prompts. The model's submitted portfolio is frozen before results are known.

  4. Step 4 Fixed wait window

    The frozen portfolio sits untouched for 7 days or 1 month.

  5. Step 5 Prices score it

    Real ending prices decide which model did best.

Benchmark universe

What can models choose from?

The active roster, asset menu, horizons, and open rounds.

Models 13
Asset choices 70
Round lengths 2
Open rounds 15
Protocol Single-turn Non-agentic calls
Latest official results

What happened in the latest scored rounds?

Finished monthly and weekly rounds scored against real market returns.

Monthly result1 of 61
Monthly official result

Monthly result scored Sep 30

Same-window returns, ranked after final prices.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
GPT-5.6 Sol
Claude Fable 5
Grok 4.6
Grok 4.3
Gemini 3.1 Pro
Claude Opus 5
Grok 4.5
S&P 500
SMH Semiconductors
GPT-5.6 Sol OpenAI
-0.34%
Claude Fable 5 Anthropic
-0.34%
Grok 4.6 xAI
-0.58%
Grok 4.3 xAI
-2.33%
Gemini 3.1 Pro Google
-3.91%
Claude Opus 5 Anthropic
-4.08%
Grok 4.5 xAI
-6.09%
S&P 500 Benchmark
-0.58%
SMH Semiconductors - Hindsight best asset
9.41%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
GPT-5.6 Sol OpenAI
Semiconductors (SMH) 35% Regional Banks (KRE) 35% Small Value (IWN) 30%
2
Claude Fable 5 Anthropic
Regional Banks (KRE) 35% Semiconductors (SMH) 35% Small Value (IWN) 30%
3
Grok 4.6 xAI
S&P 500 (SPY) 100%
4
Grok 4.3 xAI
Regional Banks (KRE) 35% S&P 500 (SPY) 65%
5
Gemini 3.1 Pro Google
Utilities (XLU) 35% Brazil (EWZ) 35% Defense (ITA) 30%
6
Claude Opus 5 Anthropic
Regional Banks (KRE) 35% Small Value (IWN) 35% S&P 500 (SPY) 30%
7
Grok 4.5 xAI
Regional Banks (KRE) 35% Small Value (IWN) 35% Real Estate (XLRE) 30%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

SMH Semiconductors - Hindsight best asset

100% Semiconductors (SMH) hindsight ceiling

Official scored round

Monthly result scored Sep 30

Audit ID: CB-2026-08-30-1M

ScoredSep 30WindowAug 31 to Sep 30Models7Asset choices70LeaderGPT-5.6 SolHorizonMonthly
Live dashboard

What is still in progress?

Open portfolios, interim returns, and upcoming score dates.

Open rounds 15 13 monthly / 2 weekly
Frozen portfolios 105 27 assets currently held
Latest close Oct 1 Live returns update before final scoring
Next score Oct 2 Official results publish after ending prices
Audit packet

How can you verify the benchmark?

Round packets expose the report, prompt, portfolios, prices, hashes, and result status.