Handcrafted games can design difficulty curves. A level designer places combat rooms in a sequence, adjusting enemy counts and room constraints based on their estimate of player fatigue and skill development. The curve is authored and the game plays it back. A procedural game cannot do this in the same way: the generator does not know what sequence of rooms the player has experienced, what their current skill level is, or how their resource state relates to what they are about to face.
This is one of the more interesting design constraints in building a procedural roguelite. We cannot author the curve. We can design the generator to produce curves that emerge from its rules. This post describes how we approach that problem and where we are still working on it.
Why Static Threat Values Are Not Sufficient
The simplest approach to difficulty scaling in a procedural dungeon is to assign threat values to enemy units and room configurations, then constrain the generator to keep cumulative threat within a defined range per depth tier. You say "depth one rooms must have threat value between 10 and 20, depth two between 25 and 40," and the generator draws from the available unit and configuration pool to stay within those bounds.
We used this approach early in development. It produces consistent baseline difficulty and it is easy to tune. It has two significant problems.
The first is that threat values are a static model of a dynamic interaction. A room with threat value 30 is more or less challenging depending on the player's health, their current ability loadout, which abilities they have recently used, and what kinds of enemies they have learned to handle efficiently. Two players at identical health with identical resources can have very different experiences of the same threat-value-30 room based on their accumulated skill with specific enemy types. A static threat value does not capture this.
The second problem is that within-run progression is invisible to the static model. A player who cleared the first four combat rooms without taking damage is in a very different state than a player who spent all their health potions reaching the same point. Both players face the same distribution of depth-two encounters because the depth-two threat band is fixed. The generator has no memory of what the run has looked like so far.
How We Approach Within-Run Scaling
We do not use dynamic difficulty adjustment in the sense of monitoring player performance and actively modifying encounter parameters in response. We considered this and decided against it because it conflicts with one of our core design commitments: runs are seeded and deterministic. If the generator's outputs depend on the player's runtime performance, the same seed can produce different dungeons for different players or for the same player on repeat attempts. That undermines the seed system's value for both reproducibility and community play.
Instead, we use a run-state factor: a scalar computed at generation time from the player's starting configuration (ability selections, difficulty mode, optional modifiers) that shifts the global difficulty band for the run. The run-state factor is itself seeded, so the same starting configuration and seed always produce the same shifted difficulty band. This preserves determinism while allowing the generator to produce runs that are systematically harder or easier based on starting conditions.
Within a run, difficulty progression is handled by the grammar's depth constraints. Each depth tier has a defined threat band (with some variance), and deeper tiers always have higher bands. This produces the expected escalation. What the grammar also enforces is the rest-distribution rule: combat rooms are interspersed with resource rooms at a minimum interval, so the player always has access to resources before the difficulty escalates significantly. The generator does not know the player's current health, but it knows that a resource room must have appeared within the last four rooms, which provides a structural guarantee rather than a dynamic one.
Difficulty Modes and Their Effect on Generation
Ruinveil has three difficulty modes: Standard, Veteran, and Cursed. The difference between these modes is not simply a stat multiplier. Each mode shifts the generation parameters in several ways: the threat band ranges per depth tier, the minimum and maximum room counts, the constraint weights that govern encounter composition, and the frequency of special room types.
Cursed mode, for example, reduces the minimum interval between resource rooms and increases the frequency of support-class enemies in the ecology system. This creates encounters that are harder not because enemies have more health but because the room-level tactical complexity is higher and recovery opportunities are less predictable. The difficulty is in the encounter design, not in the numbers.
We made this choice because we found that stat multiplication alone produces difficulty that is experienced as attrition rather than challenge. Harder enemies that take more hits are frustrating in the same way across all player skill levels. Harder encounter compositions that require different tactical approaches are more interesting for skilled players and can be learned by less-skilled players over time. We want difficulty to reward engagement with the game's systems, not just penalize players for not having high enough numbers.
The Calibration Problem
Our current calibration has a gap we want to be transparent about. The threat value system models individual enemy units and their contributions to room difficulty. It does not yet model enemy synergies as a separate difficulty multiplier. A room with an anchor and two harassers is not simply more difficult by the sum of those units' threat values: the anchor-harasser combination is specifically challenging because the harasser pressure is most damaging when the player's movement is constrained by the anchor's presence. The synergy is a multiplicative effect that our additive threat model does not capture.
We know this because our playtest data from December 2025 and March 2026 shows that anchor-harasser compositions at mid-dungeon depths produce death rates significantly above what the threat value model predicts. The rooms are harder than their threat value suggests. We are working on extending the threat model to account for pairwise role synergies, but it is a more complex model and we have not finalized the implementation.
The second calibration gap is in how difficulty modes interact with player ability selections. Some ability combinations are significantly more effective against specific ecology compositions than others. In Standard mode this is interesting: good ability-to-encounter matching feels like reward for smart selection. In Cursed mode, the gap between optimal and suboptimal ability choices for a given run's ecology is wider, which can produce runs where an ability selection that looks reasonable at the start becomes increasingly disadvantaged by depth three. This is partially intentional, but the magnitude is not well calibrated yet. We are gathering data on this from the current beta.
What We Are Not Trying to Do
We are not trying to build a game that adjusts to every player's skill level dynamically. Dynamic difficulty adjustment in roguelites has a design cost that we think outweighs the benefit: it softens the player's ability to develop genuine mastery because the game compensates for their current level. Players who improve at the game should encounter evidence of that improvement in the form of runs that go further and encounters that they handle more efficiently. If the game adjusts upward as they improve, the improvement becomes invisible to them.
What we want is a generator that produces a consistent challenge structure that players can learn over multiple runs, with enough variety in encounter composition and dungeon layout that the learning is always engaging new combinations rather than the same memorized patterns. The difficulty curves should be stable and learnable; the content within those curves should be procedurally varied. That is the balance we are working toward.