Benchmark menu

Model results

GPT 4.1 Nano

See how GPT 4.1 Nano performed, which reasoning settings were tested, how often it produced a working player program, and what those runs cost.

GPT 4.1 Nano

Provider: OpenAI

Rank
22
Status
Partial 1/3
Rank change
Baseline
Score
39.7
Last tested
2026-08-14
Settings tested
Intensive not supported; Balanced not supported

Why this result is partial: one or more bracket scores are missing; three distinct provider settings are not documented; one or more documented settings are untested or ineligible; generation-quality evidence is missing; the official generation cohort is incomplete. The rank uses the available reasoning-setting results.

Scores by reasoning setting

Available settings

Reasoning field omitted

Checked on 2026-07-31 · Provider documentation

Scores from tested settings

Tested settingScoreResult statusDisplay group
Reasoning field omitted39.7Tested and includedBaseline

How settings are grouped

  • Baseline: Reasoning field omitted — tested and included
  • Balanced: Not available for this model
  • Intensive: Not available for this model

Player program results

Reasoning field omitted: 100% / 0% / 0% · n 1 Intensive not supportedBalanced not supported

Generation cost

Average estimated cost
$0.0023
Median estimated cost
$0.0023
Estimated range
$0.0023–$0.0023
Cost data
1 combinations / 1 recorded runs
Observed total
$0.0023
Price source
Combined: Recorded attempt + Standard list-price estimate
Price list
boardgame-list-prices-2026-07-12 + boardgame-list-prices-2026-08-08
Average output
2.0k tokens
Output range
2.0k–2.0k tokens
Samples
1

Estimated cost gives equal weight to each model, game, and reasoning-setting combination.

Test setup

  • Combines this model's tested reasoning settings
  • Uses the standard GameBench player-program prompt