Civilization V is a strategy game where you lead a nation from its first village to the space age. A single game takes hundreds of turns, and the winner is often decided by choices made long before the end. That makes it a good test of whether an AI can plan ahead.
19AI models
30setups tested
813games played
411turns per game
3fixed starts
Latest news
Leading right nowGLM-5.3, deciding every 5 turns, rated 1634▶ Watch a game
Who plays best?
Each bar shows how far an AI setup's rating sits above or below the built-in AI, the dashed line.
Victory matchups: Opus-5.5-Simple-Per-5 has the highest per-player matchup victory rate: 27.1% (13/48; expected 12.5% in 8-player games.)
Shows wins per player appearance in shared games and final-score margins. The equal-chance victory rate is 1 divided by the number of players in each game (12.5% for eight players).
Module: ratings.outcome_matchups
table: panel; include_score_ratio: True; victory_rate_unit: wins per player; expected_win_rate: Equal chance: 1 / full game player count, averaged over player appearances.; p_value_win_rate: Binomial test against equal chance, available for fixed game sizes with one appearance of the player type per game.; display: vs_reference; score_ratio_margin: row minus column
Opus-5.5 (every 5 turns) won 27% of its games. With 8 players, an even chance would be 1 in 8.
Prediction quality: attention performs best on roc auc at 0.8545 using 770 games; scores range from 0.8066 to 0.8545.
Measures how well each estimator identifies likely winners and matches observed outcomes (discrimination and calibration).
Module: prediction.evaluate
metrics: roc_auc, brier_score, log_loss, balanced_accuracy; n_models: 3
Shown a winner and a loser from the same game, our best predictor picks the winner 85% of the time, well before the game ends.
Usage, cost, and skill: Most cost-efficient: GLM-5.3-Simple-Per-5 (+99 Elo vs the fitted curve). Least cost-efficient: Qwen-3.6-27B-Simple-Per-5 (-116 Elo vs the fitted curve).
Compares cost and token use per player per game with skill, and measures Elo above or below the fitted usage-skill curve.
Module: performance.usage_efficiency
currency: usd; log_x: True; ratings_stage: bt_main; dropped_baselines: 1; unpriced_identities: 0; unrated_identities: 0; cost_basis: per player per complete game; cached_input_estimated: True; efficiency_metric: elo - expected_elo; usage_skill_equation: expected_elo = intercept + slope * log10(average_usage); usage_skill_fits: {'cost': {'intercept': 1441.9160969832785, 'slope': 70.82723827641072, 'n': 30, 'r_squared': 0.4018791576026215}, 'input': {'intercept': 929.1500568077244, 'slope': 78.62554613635585, 'n': 30, 'r_squared': 0.03811646766643484}, 'output': {'intercept': 1221.2790332535012, 'slope': 43.50647797457705, 'n': 30, 'r_squared': 0.10198755780836999}}; baseline_elo: 1500.0; baseline_name: Vanilla; null_baseline_elo: 1311.604036752814
GLM-5.3 (every 5 turns) gets the most skill for its price. Qwen-3.6-27B (every 5 turns) gets the least.
Strategic settings: Against the completed-experiment average, the largest departure is Qwen-3.6-27B-Simple | Every-turn on waterconnection (-31).
Shows how each strategist sets the in-game AI's flavors (0 to 100, 50 is balanced), against the average of completed experiments on the same map and seat and in absolute terms.
Module: behavior.flavors
n_absolute_players: 1444; baseline: completed-experiment average; n_relative_players: 1444; n_unmatched_controlled_players: 0; n_baseline_players: 1444; n_baseline_experiments: 30; flavors: Offense, Defense, CityDefense, Mobilization, MilitaryTraining, Recon, Ranged, Mobile, Nuke, UseNuke, Naval, NavalRecon, Air, Antiair, AirCarrier, Airlift, Expansion, Growth, TileImprovement, Infrastructure, Production, Gold, Science, Culture, Happiness, NavalGrowth, NavalTileImprovement, WaterConnection, GreatPeople, Wonder, Religion, Diplomacy, Espionage, Spaceship
Against the completed-experiment average, the largest departure is Qwen-3.6-27B (every turn) on waterconnection (-31).
Policy paths: Freedom: Kimi-K2.7-Simple | Every-turn (29%) · Autocracy: Qwen-3.6-27B-Simple | Every-turn (33%) · Order: GLM-5.3-Flash-Simple | Per-5 / Kimi-K2.7-Simple | Every-turn / Qwen-3.5-Simple | Per-5 / Qwen-3.8-27B-Simple | Every-turn (58%)
Shows which policy branches and ideologies each player type adopts, which it picks first in each tier, and how early, against the in-game AI on the same map and seat.
Module: behavior.policies
n_absolute_players: 1492; baseline: matched in-game AI; baseline_experiments: vanilla-standard-fixed; n_relative_players: 1492; n_unmatched_controlled_players: 0; n_baseline_players: 192; branches: tradition, authority, progress, fealty, statecraft, artistry, industry, imperialism, rationalism, freedom, autocracy, order
@article{chen2026civbench,
title={CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V},
author={Chen, John and Cheng, Sihan and Gurkan, Can and Lin, Mingyi},
journal={Conference on Language Modeling (COLM)},
year={2026}
}
Benchmark results
@misc{civbench_results,
title={Controlled CivBench on Vox Populi 5.2.7, Civilization V},
url={https://github.com/vox-deorum/civ-bench}
}