
With its strong addictiveness and immersive quality, "Sid Meier's Civilization VI" (hereinafter Civ VI), known as one of the three devil's games, saw top-tier AIs collapse helplessly. They showed insufficient stamina in long matches lasting up to eight hours. This explains why recent global AI companies have identified strengthening 'long-term task execution capability' as a core priority.
A joint research team including Oxford University in the UK recently published a paper titled "CivBench: A Benchmark for Evaluating Long-Term Task Execution of Agents Using Tools in Civilization VI" on the preprint site arXiv earlier this month. The paper analyzes results from testing four major AI models in Civ VI and presents them as a benchmark.
The "Civilization" series is a turn-based strategy game where players lead a nation from ancient times to the future, competing against other nations. Each game involves over 300 turns (rounds), requiring thousands of decisions while considering various variables such as economy, military, science, diplomacy, and culture—making it an ideal research tool. The four major AI models tested were: Anthropic's "Claude Opus 4.6," OpenAI's "GPT-5.4," Google's "Gemini 3.1 Pro," and Moonshot AI's "Kimi K2.5."
The AIs received a poor performance record of "3 wins, 20 losses" across 23 total games against built-in computer opponents. They won three out of 19 games at the Prince (normal) difficulty level but failed to win any of the four games at the higher King difficulty level. Specifically: Opus 4.6 achieved 2 wins and 6 losses; Gemini 3.1 Pro recorded 1 win and 5 losses; GPT-5.4 had 0 wins and 8 losses; Kimi K2.5 had 0 wins and 1 loss. At King difficulty, both Opus 4.6 and GPT-5.4 played two games each. However, the research team stated, "Given the small sample size, it is difficult to definitively rank the models against one another."

The main flaw was "failure to monitor the situation." The research team instructed the AIs to check their progress every 20 turns, but in practice, the AIs only reviewed the situation once every 30 to 75 turns. The researchers noted, "Monitoring frequency did not increase even in the later stages where situational checks were most useful," adding that this indicates a common trait inherent to AI models rather than temporary errors.
Execution capability was also lacking. Although the AIs formulated plans such as "building a science campus in a new city" or "constructing a second city," they failed to execute them within 10 turns. A "Plan Execution Score" was calculated on a scale of 100 points: 1 point for completing a pre-set plan within 10 turns, 0.5 points for partial execution, and 0 points for non-execution. The AIs scored only between 48.2 and 65.8 points. The researchers explained, "Improving reasoning ability alone is insufficient to ensure reliability and completeness in long-term tasks. Measures such as mandating regular status checks or prioritizing monitoring tools are necessary."
In fact, global AI companies are also striving to go beyond single-shot responses and develop the capability to complete goals fully. Kim Kyung-hoon, Head of OpenAI Korea, stated on the 9th regarding their latest model "GPT-6 Astra," "We focused development on whether tasks can be completed from start to finish. The core capability lies in continuously recognizing goals, remembering prior task context, and delivering high-quality results."
Anthropic also explained at the early release of "Claude Opus 4.6" earlier this year that the model is designed to carefully formulate long-term plans and sustain work over extended periods. Anthropic has even introduced a feature that automatically summarizes prior context to enable continuous long-term tasks. Google similarly emphasized in May that Gemini 3.5 Flash is well-suited for long-term agent operations.
