AI Opponent Logic for Creature Battle Games
Balancing fairness with challenge when the AI can't see what players are hiding.

Creature battle AI has to do two things that fight each other. It needs to feel threatening and adaptive, but never unfair. And it needs to handle information the player is actively hiding from it. Most combat AI in games, by contrast, gets to see everything on screen and just has to react fast enough. Game AI design research lays out the core conflict: designers want believable opponents that respond sensibly to what the player does, but they also want to shape specific experiences that push players toward new strategies. Those two goals don't sit comfortably together. An opponent tuned purely to react "sensibly" can feel flat. One tuned to surprise the player can feel like it's cheating.
Creature battles add a wrinkle most combat games skip. Each side picks a move without seeing the other side's choice, so the AI is working with real partial observability. It doesn't know the opponent's full team, their held items, or how they built their stats. It can only guess. Some systems guess using meta-statistics, pulling from what's common at a given skill level. Others run combinatorial optimization, like mixed-integer programming, to generate plausible rosters the opponent might be running. Either way, the AI has quietly swapped problems. It's answering "what's the best move, given a spread of things the opponent might have, each with some probability attached.
Stack that uncertainty on top of combinatorial move, item, and matchup possibilities, and add a hard time limit per turn, and brute-force search runs out of road fast, long before anyone even starts thinking about making the opponent behave in an interesting way. The PokéAgent Challenge, held at NeurIPS 2025, picked Pokémon battles as a testbed for this kind of problem. The organizers framed it as sitting at the intersection of reinforcement learning, game theory, planning, and language models, because the genre surfaces weaknesses that simpler combat setups let slide. Any architecture built for creature battles has to answer to that partial observability and that combinatorial load before it earns the right to be clever.
Finite state machines and the basic skeleton of creature AI
Finite state machines are where almost every creature AI starts, and for good reason. They give a creature the bare minimum it needs to look like it's making decisions. A creature that knows when to patrol, chase, or flee is already doing something a coin flip couldn't.
A typical FSM-driven battle creature runs on three states. It patrols when no player is around. It switches to chasing and attacking once a player shows up. It flips to escaping once its health drops below some set line. Each state follows the same basic shape: entry logic for what happens the moment the state kicks in, execution logic for what runs every tick while it's active, and exit logic for what triggers the jump to the next state.
The appeal is that this is predictable. Predictable means consistent, and consistent means a designer can actually test and debug it without pulling their hair out. But that same predictability is the problem. A player who's paid attention for ten minutes can read the transitions and route around them, the same way you'd learn a guard's patrol pattern in a stealth game. Pile on more states to cover more situations, and the number of transitions between them grows combinatorially. Developers call this transition explosion, and it turns a tidy diagram into a tangle that breaks the moment someone tries to add one more behavior. None of this makes FSMs worthless. They're still a solid way to manage the big-picture buckets, like whether a creature is in combat, exploring, or socializing. They're still just not enough, by themselves, to make a creature feel like a real threat.
What behavior trees add that state machines cannot provide
Behavior trees exist because someone got tired of drawing transition diagrams that looked like subway maps. They became the standard way to handle complex NPC decisions because they sidestep the FSM's transition problem entirely, organizing logic into a tree that's modular, reusable, and checked from the top down on every tick.
Instead of wiring up an explicit transition between every possible pair of states, a behavior tree checks conditions from the root down to the leaves each tick, and runs whichever branch's conditions are met first. Priority is baked into the shape of the tree itself, not into some separate web of arrows. That alone makes the whole thing modular. A subtree built to handle flanking, or retreating, or using a particular category of move, can be built and tested on its own, then slotted into something bigger.
A lot of production games split the work so the FSM handles the macro question of whether this creature is in combat right now, while the behavior tree handles everything underneath that: which move to use, which target to pick, when to change tactics mid-fight. That division is what lets creature AI scale across a whole roster. A new creature can reuse the common subtrees everyone shares and just bolt on a few leaves for whatever unique ability makes it special, instead of someone rebuilding the entire decision system from scratch every time.
What behavior trees don't fix is the rule-based bones underneath. The priority order is still something a designer set in advance, at build time. The tree adapts to what's happening in the game, sure, but it doesn't adapt to the specific player sitting across from it. A patient, observant player can still map out the fixed priority order and play around it, the same way they would with an FSM, just with more steps involved.
How utility-based and goal-oriented systems weigh options
Utility-based systems change the question the AI is answering. Instead of "what's the first valid option on the list," it becomes "what's the best option given everything going on right now." That shift alone makes the resulting decisions harder to predict and better suited to the actual moment.
In a utility-based setup, every possible action gets scored against the current state of the battle. A heavy-damage move scores higher when the opponent's health is already low. A defensive move scores higher when the creature's own health is dropping. The AI just takes whichever option scores best, rather than grabbing the first branch that happens to qualify. Goal-oriented action planning, or GOAP, pushes this idea a step further: instead of scoring individual actions, the AI sets a target end state, like forcing the opponent's creature to faint or keeping its own HP above some floor, and works backward to find a sequence of moves that gets there.
Both approaches let the AI respond to differences in degree, not just differences in category. A creature leans more and more toward defensive options as the health bar drops, the way a person gets more cautious the closer they get to the edge of a cliff. That matters a lot in creature battles specifically, because the variables in play (HP, type matchups, status effects, moves left) are plentiful and they interact in ways a fixed-priority tree just can't represent cleanly.
None of this comes free. Scoring every option and planning backward from a goal costs more compute than just walking down a tree. The output is only as good as the math behind it. Badly calibrated utility functions produce an AI that technically weighs everything and still makes decisions that make players tilt their heads sideways. Gladiabots, a robot combat strategy game, gives a good sense of how far conditional action selection can go even without a dedicated goal-planning system. Players build their squad's logic with a no-programming system that offers millions of possible combinations, and the resulting behavior on the battlefield gets surprisingly sophisticated. It's proof that this kind of layered decision-making is legible enough for regular players to build with their own hands, not just something locked inside an engineer's codebase.
How reinforcement learning changes what the AI can learn about an opponent
Reinforcement learning hands creature AI something none of the earlier layers can: strategies nobody explicitly wrote down. The AI develops them through trial and error against simulated opponents. Once training ends, the policy is locked. It doesn't keep adjusting to the particular player on the other side of the screen.
Different projects lean on different flavors of this. Classic reinforcement learning is one path. PokéChamp pairs minimax search with language model modules. PokéLLMon combines in-context reinforcement learning with knowledge-augmented generation. Each one targets a different weak spot, whether that's hallucination, modeling what the opponent is likely holding, or refining the policy over time. What all of them share is the same tradeoff: the learned behavior can be devastating against the range of opponents it trained against, but it's still a fixed policy once it ships.
PokéLLMon shows what that training can actually buy you. Tested against ladder players on Pokémon Showdown, it performed comparably to disciplined, experienced ladder players, which is a real signal that RL-trained creature battle AI can hit a genuinely competitive level, not just a passable one. The PokéAgent Challenge at NeurIPS 2025 built standardized benchmarks specifically around opponent modeling under partial observability and long-horizon reasoning, because creature battles expose RL weaknesses that simpler game environments never bother to test.
Deployment is the weak point. Once an RL agent is out in the world, it can't readily adapt. If a player finds an exploit, or just plays in a style the training data never covered, the agent has no built-in way to update in response. Scripted opponents are predictable and RL opponents are adaptive only during training, then static afterward, so neither option has fully satisfied the industry, which has pushed developers toward hybrid designs. There's a telling data point from D&D 5E combat experiments: in a four-class tournament, RL models took the top spots, including one trained against rules-based opponents as well as other language models. In a fighter-only tournament, language model agents took the top spots instead. The pattern suggests RL performs best when it trains against a genuinely diverse, high-quality field of opponents, not against endless copies of itself.
What LLM augmentation adds to reinforcement-learned creature AI
Pairing a trained RL agent with a language model lets that agent shift its tactical posture on the fly, reacting to live game state instead of sticking to whatever it learned in training, which is the pitch. The evidence so far says the language model's real-time strategic judgment is a lot shakier than the pitch suggests.
A 2026 study trained NPC agents on a shared PPO policy inside Unity and compared a plain baseline against a version augmented by a locally hosted Mistral model. That model read the live game state every five seconds and assigned one of four tactical tags to guide the agent. Against an aggressive opponent, the language model kept recommending encirclement almost every time, and that habit actively hurt performance in exactly the matchup where a different approach was needed. Digging into the tag choices across the whole study found the same pattern everywhere: the model picked "Surround" overwhelmingly, regardless of who it was facing. It wasn't reading the matchup and adjusting. It was defaulting to one answer and calling it strategy.
The study's own authors state that the results "characterize both the promise and the current ceiling of language-model-guided runtime strategy selection as a design pattern for adaptive game AI. That ceiling is acknowledged in the research itself, not an outside complaint. The sharpest objection to this whole approach is that any performance gains might come from giving the squad one consistent signal to coordinate around, not from the model actually reasoning its way to a smart call. Under that reading, the LLM is acting as a glorified metronome for squad coordination, not a battlefield strategist.
A plan to move the language model out of the moment-to-moment loop entirely is one alternative worth watching. It's a two-tier setup sometimes called TTA. The first tier trains a batch of diverse, skilled deep RL agents. The second tier uses an LLM, sometimes called a Hyper-Agent, that looks at player data and feedback and picks which of those pre-trained agents to send into the match. The language model isn't deciding what move to make mid-fight; it's deciding who gets sent into the ring in the first place, working more like a matchmaker than a tactician.
How the layers interact in practice
None of these layers is doing anything remarkable on its own. The reason the full system works is that each layer handles the kind of reasoning it's actually good at, then passes the baton cleanly to the next one.
In practice, the FSM decides whether the creature is in combat at all, while the behavior tree picks which category of move is even worth considering once combat starts. Utility scoring or GOAP picks the specific action that wins within that category. And an RL policy, or an LLM layer sitting above it, nudges the overall tactical posture over a longer stretch of the fight. Scorched Sun, released September 25, 2026 by developer Skeleton Coffee Games for PC, shows what this looks like from the player's side of the screen. The game lets players target individual body parts, and severing a creature's limb can change how that creature behaves, often making it more dangerous. That's a player staring down a direct, visible consequence of state-responsive design: cutting off a limb isn't just damage, it's a state transition, and a smart player has to think about what that transition might unlock before they commit to the strike.
The Pokémon Company's TCG AI Battle Challenge, with its Final Stage scheduled for September 2026 in Japan and matches streamed on Pokémon's official YouTube channel, gives a sense of how much a single architecture is expected to cover. The competition requires agents that read an opponent's deck, hand, board state, likely strategies, and card combinations, all while the match is actually happening. The format alone tells you how wide the net has to be cast.
Where this whole layered approach falls apart is at the seams, when the layers are built in isolation and nobody specifies cleanly how they hand off to each other. An AI can pick a genuinely sharp individual move and still get read like a book if its shift between macro states is obvious. A creature that telegraphs the moment it's about to switch from aggressive to defensive is still giving the game away, no matter how good its move selection is underneath.
What makes an AI opponent feel genuinely threatening
An opponent feels like a real threat when its decisions are built around the specific player it's facing, not some average player the designers imagined. That's less a question of how fancy the architecture is and more a question of dynamic difficulty and opponent modeling done well.
Dynamic AI systems track in-game signals as the match unfolds: how the player is performing, how much health they have left, the patterns in the choices they keep making. The system uses those signals to adjust the challenge on the fly, which might mean dialing the AI's aggression up or down, shifting how resources get distributed, or tweaking other parameters mid-match. The goal is a challenge that's shaped around the player actually in the seat, instead of a difficulty slider set once and left alone for everyone. Every layer discussed here, the states, the trees, the utility scores, the learned policies, the language model sitting on top, exists to serve that one outcome: an opponent that reads the player in front of it and responds like it means it.
Sources
- GLADIABOTS - AI Combat Arena - Apps on Google Play
- A Resourceful Reframing of Behavior Trees
- AI for Games in the Foundation Model Era
- Toward Smarter Opponents: Rethinking AI in Turn-Based Games — Arts Management and Technology Lab
- Development of Two Dimension (2D) Game Engine with Finite State Machine (FSM) Based Artificial Intelligence (AI) Subsystem - ScienceDirect
- Augmenting Game AI with Deep Reinforcement Learning
- AI - State Machines Introduction
- Component-based hierarchical state machine — A reusable and flexible game AI technology
