CivBench · Civilization V with Vox Populi 5.2.7

Which AI makes the best long-term strategist?

Civilization V is a strategy game where you lead a nation from its first village to the space age. A single game takes hundreds of turns, and the winner is often decided by choices made long before the end. That makes it a good test of whether an AI can plan ahead.

Latest news

Leading right nowGLM-5.3, deciding every 5 turns, rated 1634 ▶ Watch a game

Who plays best?

Each bar shows how far an AI setup's rating sits above or below the built-in AI, the dashed line.

Show the other 22

What is this built on?

Three pieces of software make these games possible. You do not need to play any of them to read this report.

Civilization V

A 2010 strategy game by Firaxis. You lead a nation through history one turn at a time: building cities, trading, making friends and enemies.

Steam store page ↗

Vox Populi

A large community mod that rebalances Civilization V and makes its computer players much smarter. All games here use version 5.2.7.

Vox Populi on GitHub ↗

Vox Deorum

Our open-source bridge that lets an AI agent read the game and steer a nation (through the execution of the built-in AI).

Vox Deorum on GitHub ↗

Watch any game in the replay viewer ↗ · Read the CivBench paper ↗ · Read the Vox Deorum paper ↗

Is the pricier model worth it?

More compute usually buys a bit more skill. Setups above the dashed line play better than their price suggests.

  • every turn
  • every 5 turns
25¢ 50¢ $1 $2 $5 $10 $20 1350 1400 1450 1500 1550 1600 1650 Built-in AI Typical rating for the price Cost per player per game, US dollars (each step multiplies the cost) Rating GLM-5.3 Kimi-K2.7 Opus-5.5 Qwen-3.6-27B Opus-5.5 DeepSeek-V4-Flash DeepSeek-V4-Flash DeepSeek-V4.1-Flash GLM-5.1 GLM-5.2 GLM-5.2 GLM-5.3-Flash GLM-5.3-Flash GLM-5.3 GPT-6-Luna GPT-OSS-120B GPT-OSS-120B Gemma-4 Gemma-4 Kimi-K2.6 Kimi-K2.7 Kimi-K2.7 MiniMax-M2.7 MiniMax-M2.7 MiniMax-M3 Nemotron-3-Super Nemotron-3-Super Qwen-3.5 Qwen-3.5 Qwen-3.6-27B Qwen-3.6-27B Qwen-3.8-27B Qwen-3.8-27B Qwen-3.8-Flash-Next

How do they like to win?

Each bar shows how much of the game an AI spent aiming for each victory type.

Key findings

One sentence each. The ? button shows the technical version, and the link opens the full results.

Strongest player

Pairwise skill ratings: GLM-5.3-Simple-Per-5 leads 32 identities at 1634 Elo; rating spread: 322 points. Estimates each player type's relative skill from pairwise comparisons of model-adjusted strength within each game (Bradley-Terry Elo ratings). Module: ratings.bradley_terry group_by: player_type; strength_table: strength; adjust_stage: strength; strength_estimator: attention; estimator_model: attention_mlp; estimator_fit: pretrained; estimator_predict: in_sample; adjust_block: auto/start_cell; adjust_turn_progress_min: 0.2; adjust_weight: turn_progress; adjust_enforce_winner: True; adjust_civ_adjust: ols_logit; adjust_baseline_experiment: vanilla-standard-fixed; adjust_post_cell_normalize: none; adjust_cell_gain_bend: 0.5

GLM-5.3 (every 5 turns) is the strongest so far. It beats the built-in AI 68% of the time head-to-head.

See details →

Most wins

Victory matchups: Opus-5.5-Simple-Per-5 has the highest per-player matchup victory rate: 27.1% (13/48; expected 12.5% in 8-player games.) Shows wins per player appearance in shared games and final-score margins. The equal-chance victory rate is 1 divided by the number of players in each game (12.5% for eight players). Module: ratings.outcome_matchups table: panel; include_score_ratio: True; victory_rate_unit: wins per player; expected_win_rate: Equal chance: 1 / full game player count, averaged over player appearances.; p_value_win_rate: Binomial test against equal chance, available for fixed game sizes with one appearance of the player type per game.; display: vs_reference; score_ratio_margin: row minus column

Opus-5.5 (every 5 turns) won 27% of its games. With 8 players, an even chance would be 1 in 8.

See details →

How reliable is this?

Prediction quality: attention performs best on roc auc at 0.8545 using 770 games; scores range from 0.8066 to 0.8545. Measures how well each estimator identifies likely winners and matches observed outcomes (discrimination and calibration). Module: prediction.evaluate metrics: roc_auc, brier_score, log_loss, balanced_accuracy; n_models: 3

Shown a winner and a loser from the same game, our best predictor picks the winner 85% of the time, well before the game ends.

See details →

Best value

Usage, cost, and skill: Most cost-efficient: GLM-5.3-Simple-Per-5 (+99 Elo vs the fitted curve). Least cost-efficient: Qwen-3.6-27B-Simple-Per-5 (-116 Elo vs the fitted curve). Compares cost and token use per player per game with skill, and measures Elo above or below the fitted usage-skill curve. Module: performance.usage_efficiency currency: usd; log_x: True; ratings_stage: bt_main; dropped_baselines: 1; unpriced_identities: 0; unrated_identities: 0; cost_basis: per player per complete game; cached_input_estimated: True; efficiency_metric: elo - expected_elo; usage_skill_equation: expected_elo = intercept + slope * log10(average_usage); usage_skill_fits: {'cost': {'intercept': 1441.9160969832785, 'slope': 70.82723827641072, 'n': 30, 'r_squared': 0.4018791576026215}, 'input': {'intercept': 929.1500568077244, 'slope': 78.62554613635585, 'n': 30, 'r_squared': 0.03811646766643484}, 'output': {'intercept': 1221.2790332535012, 'slope': 43.50647797457705, 'n': 30, 'r_squared': 0.10198755780836999}}; baseline_elo: 1500.0; baseline_name: Vanilla; null_baseline_elo: 1311.604036752814

GLM-5.3 (every 5 turns) gets the most skill for its price. Qwen-3.6-27B (every 5 turns) gets the least.

See details →

Changed habits

Strategic settings: Against the completed-experiment average, the largest departure is Qwen-3.6-27B-Simple | Every-turn on waterconnection (-31). Shows how each strategist sets the in-game AI's flavors (0 to 100, 50 is balanced), against the average of completed experiments on the same map and seat and in absolute terms. Module: behavior.flavors n_absolute_players: 1444; baseline: completed-experiment average; n_relative_players: 1444; n_unmatched_controlled_players: 0; n_baseline_players: 1444; n_baseline_experiments: 30; flavors: Offense, Defense, CityDefense, Mobilization, MilitaryTraining, Recon, Ranged, Mobile, Nuke, UseNuke, Naval, NavalRecon, Air, Antiair, AirCarrier, Airlift, Expansion, Growth, TileImprovement, Infrastructure, Production, Gold, Science, Culture, Happiness, NavalGrowth, NavalTileImprovement, WaterConnection, GreatPeople, Wonder, Religion, Diplomacy, Espionage, Spaceship

Against the completed-experiment average, the largest departure is Qwen-3.6-27B (every turn) on waterconnection (-31).

See details →

Diplomacy

Diplomatic behavior: Friendliest: Nemotron-3-Super-Simple | Every-turn (+207.4 net) · Least friendly: MiniMax-M3-Simple | Per-5 (-37.3 net) · Most masked: DeepSeek-V4-Flash-Simple | Every-turn (4.4%) Describes diplomatic persona traits and the public and private stances strategists set toward rivals, including how often the two conflict. Module: behavior.diplomacy n_absolute_players: 1444; baseline: matched in-game AI; baseline_experiments: vanilla-standard-fixed; n_relative_players: 1444; n_unmatched_controlled_players: 0; n_baseline_players: 192; traits: DiplomaticBalance, Friendliness, WorkWithWillingness, WorkAgainstWillingness, Loyalty, DenounceWillingness, Forgiveness, Meanness, Neediness, Chattiness, DeceptiveBias; rate: per_100_turns

Friendliest: Nemotron-3-Super (every turn) (+207.4 net) · Least friendly: MiniMax-M3 (every 5 turns) (-37.3 net) · Most masked: DeepSeek-V4-Flash (every turn) (4.4%)

See details →

Ways to win

Strategic commitment: Domination: Gemma-4-Simple | Every-turn (43%) · Culture: MiniMax-M2.7-Simple | Every-turn (53%) · Diplomatic: GPT-OSS-120B-Simple | Every-turn (51%) · Science: GPT-6-Luna-Simple | Per-5 (76%) Shows how often strategists act and revise their settings, how large and how lasting their changes are, and which grand strategy they hold. Module: behavior.commitment n_absolute_players: 1444; baseline: completed-experiment average; n_relative_players: 1444; n_unmatched_controlled_players: 0; n_baseline_players: 1444; n_baseline_experiments: 30; grand_strategies: Conquest, Culture, UnitedNations, Spaceship; rate: per_100_turns

Who aims for each kind of win most often. Domination: Gemma-4 (every turn), 43%; Culture: MiniMax-M2.7 (every turn), 53%; Diplomacy: GPT-OSS-120B (every turn), 51%; Science: GPT-6-Luna (every 5 turns), 76%.

See details →

Politics

Policy paths: Freedom: Kimi-K2.7-Simple | Every-turn (29%) · Autocracy: Qwen-3.6-27B-Simple | Every-turn (33%) · Order: GLM-5.3-Flash-Simple | Per-5 / Kimi-K2.7-Simple | Every-turn / Qwen-3.5-Simple | Per-5 / Qwen-3.8-27B-Simple | Every-turn (58%) Shows which policy branches and ideologies each player type adopts, which it picks first in each tier, and how early, against the in-game AI on the same map and seat. Module: behavior.policies n_absolute_players: 1492; baseline: matched in-game AI; baseline_experiments: vanilla-standard-fixed; n_relative_players: 1492; n_unmatched_controlled_players: 0; n_baseline_players: 192; branches: tradition, authority, progress, fealty, statecraft, artistry, industry, imperialism, rationalism, freedom, autocracy, order

Freedom: Kimi-K2.7 (every turn) (29%) · Autocracy: Qwen-3.6-27B (every turn) (33%) · Order: GLM-5.3-Flash (every 5 turns) / Kimi-K2.7 (every turn) / Qwen-3.5 (every 5 turns) / Qwen-3.8-27B (every turn) (58%)

See details →

Citation

Paper

@article{chen2026civbench,
  title={CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V},
  author={Chen, John and Cheng, Sihan and Gurkan, Can and Lin, Mingyi},
  journal={Conference on Language Modeling (COLM)},
  year={2026}
}

Benchmark results

@misc{civbench_results,
  title={Controlled CivBench on Vox Populi 5.2.7, Civilization V},
  url={https://github.com/vox-deorum/civ-bench}
}