Skip to content

Ruggedness

Ruggedness is arguably the most widely studied topographical aspect of fitness landscapes. Although biologists have offered a staggering number of definitions of ruggedness, they agree broadly on the essence of the idea: the lack of correlation in fitness between genotypes. This can either manifest as the presence of multiple fitness peaks (local optima) or significant fitness fluctuation.

Relationship to landscape navigability: Ruggedness can pose a fundamental challenge to evolution’s ability to find a landscape’s highest peaks, i.e., landscape navigability or peak accessibility. This is because a population evolving under the influence of natural selection can only travel on accessible paths through the landscape, that is, paths in which each mutational step increases fitness. The reason is that natural selection favors high fitness genotypes and does not allow a population to traverse low fitness valleys between a local peak of intermediate fitness and nearby higher fitness peaks.

  • In a smooth landscpae, all mutational paths to the single peak are accessible.
  • When landscape becomes rugged, many peaks exist, yet high-fitness peaks are reachable by abundant short accessible paths, especially in high-dimensional sequence spaces.
  • For a maximumally rugged landscape, selectively accessible paths become rare, and evolution may stall on suboptimal peaks.

Relationship to epistasis: Reciprocal sign epistasis is a necessary yet not sufficient condition for landscape ruggedness-landscape can be rugged even when it is rare, while its mere presence does not guarantee ruggedness. Yet, for a nonepistatic (i.e., purely additive) landscape, there is often only a single, global fitness peak. In contrast, pervasive sign epistasis could create a rugged landscape with numerous sub-optimal peaks.

Following the key idea of ruggedness, various established quantitative measures have been developed.

Overview

API Purpose
local_optima_ratio Independent local optima divided by retained configurations.
r_s_ratio Roughness-to-slope ratio — deviation from a purely additive fit.
autocorrelation Fitness autocorrelation along random walks (Weinberger 1990).
gradient_intensity Mean absolute fitness step per edge, normalized by mean fitness.
neighbor_fitness_correlation Correlation between node fitness and mean neighbor fitness.

Local Optima

Calculates the proportion of local optima variants in the landscape.

The number of local optima (peaks) is a principal indicator of fitness landscape ruggedness. To make it a unitless measure, we divide it by the number of total variants in the landscape. To provide a sense of its magnitude, a mostly rugged NK landscape with dimension \(n\) and degree of interaction \(k=n-1\) would have \(\frac{2^n}{n+1}\) local optima. When \(n=20\), the ratio of local optima would be around \(4.76\%\).

graphfla.analysis.local_optima_ratio(landscape) -> float

Return the number of local-optimum plateaus per configuration.

Parameters

landscape : Landscape

Built fitness landscape.

Returns

ratio : float

landscape.n_lo / landscape.n_configs, or NaN on an empty landscape. The numerator counts local-optimum plateaus, whereas the denominator counts configurations. A multi-member peak plateau counts once.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import local_optima_ratio
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> local_optima_ratio(landscape)
0.25

Roughness-to-Slope (r/s) Ratio

The roughness-to-slope (r/s) ratio measures how well a landscape can be described by an additive model, where individual variables contribute independently to the objective value. It is calculated by fitting a linear model to the encoded configurations using least squares. The slope is the mean absolute additive coefficient, while the roughness is the root mean square of the residuals.

A smaller ratio indicates less variation around the additive trend, reaching zero for an exact additive fit with nonzero slope. A larger ratio indicates greater residual variation relative to that trend. Comparisons require consistent variable encoding: categorical reference states can affect the ratio, and nonlinear effects of an ordinal variable can contribute to roughness even without interactions between variables.

graphfla.analysis.r_s_ratio(landscape) -> float

Return the roughness-to-slope ratio of an additive least-squares fit.

Parameters

landscape : Landscape

Built landscape. Each retained configuration has equal weight; its stored fitness is the objective value, with no log transformation. Boolean variables use 0/1 coding. Categorical variables use one indicator per observed state except the first in pandas category order (normally sorted labels). Ordinal variables use the integer ranks established during construction, or pandas category order if construction codes are unavailable. Constant variables are excluded.

Returns

ratio : float

Residual root mean square divided by the mean absolute coefficient, excluding the intercept. The mean weights encoded columns equally, not original variables. Returns zero, up to rounding error, for an exact additive fit with nonzero slope. Returns numpy.inf when slope is at most 1e-12 times the objective range. Returns numpy.nan for empty or constant data, an unidentifiable fit, a failed numerical solve, or a model exceeding the workspace limit.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import r_s_ratio
>>> landscape = BooleanLandscape().build_from_data(
...     [[0, 0], [0, 1], [1, 0], [1, 1]], [0, 1, 2, 4], verbose=False)
>>> round(r_s_ratio(landscape), 3)
0.125

Notes

With an intercept, fit \(\hat f_i=b_0+\sum_j b_j x_{ij}\) by ordinary least squares, then calculate \(r=\sqrt{\sum_i(f_i-\hat f_i)^2/N}\) and \(s=\sum_j|b_j|/p\) 1. Missing configurations are not imputed; graph edges and the optimization direction do not enter the fit.

For categorical variables with more than two states, changing the reference state can change s and the ratio. For ordinal variables, nonlinear effects of a single variable also contribute to r. Compare values only under consistent coding and objective scales.

References

[1] I. G. Szendro et al., "Quantitative analyses of empirical fitness landscapes," J. Stat. Mech. P01005 (2013), Eqs. (3)-(5). https://doi.org/10.1088/1742-5468/2013/01/P01005

Raises

graphfla.exceptions.NotBuiltError

If the landscape has not been built.

ValueError

If a variable type is unsupported or an objective value is nonfinite.

Warns

UserWarning

If the ratio is undefined, the slope is numerically zero, or the fit is saturated (no residual degrees of freedom). A fit also returns NaN with a warning if its quadratic QR workspace exceeds 64 MiB.

Autocorrelation

Calculates the autocorrelation of a fitness landscape (Weinberger 1990).

Autocorrelation measures landscape ruggedness by simulating random walks across the landscape. For a smooth landscape, fitness values for adjacent variants encountered during the same walk would be highly correlated (autocorrelation close to 1). In contrasts, this correlation diminishes in rugged landscapes wherein fitness values fluctuates dramatically even across a single mutation (autocorrelation close to 0).

graphfla.analysis.autocorrelation(
    landscape,
    walk_length: int = 20,
    walk_times: int = 1000,
    lag: int = 1,
    seed: Optional[int] = None,
) -> float

Return the pooled fitness autocorrelation along random walks.

Parameters

landscape : Landscape

Built fitness landscape.

walk_length : int, default=20

Maximum number of visited configurations, including the starting node.

walk_times : int, default=1000

Number of random walks, each starting at a uniformly sampled node.

lag : int, default=1

Separation between fitness observations along a walk.

seed : int or None, default=None

Seed for local random walks. An integer makes the sample reproducible; None uses the global Python random state.

Returns

correlation : float

Lagged fitness correlation pooled under one grand mean. Returns NaN if no walk has more than lag observations or pooled variance is zero.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import autocorrelation
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> round(autocorrelation(landscape, walk_times=8, seed=0), 3)
-0.082

References

[1] E. Weinberger, "Correlated and Uncorrelated Fitness Landscapes and How to Tell the Difference", Biol. Cybern. 63, 325-336 (1990).

Gradient Intensity

Calculates the gradient intensity of the landscape: the average absolute fitness difference (delta_fit) along all edges of the landscape graph, normalized by the mean fitness so that the result is scale-invariant.

A landscape with steep "uphill" edges relative to the typical fitness magnitude will have a high gradient intensity. A landscape whose fitness values are nearly flat (or whose typical fitness step is small compared to the overall fitness scale) will have a low gradient intensity. Together with autocorrelation and r_s_ratio, this provides a complementary numeric signature of how locally varied the landscape is.

graphfla.analysis.gradient_intensity(landscape) -> float

Return the mean absolute edge fitness change divided by mean fitness.

Parameters

landscape : Landscape

Built fitness landscape.

Returns

intensity : float

Mean absolute delta_fit over graph edges, divided by mean fitness. Missing edge delta_fit attributes contribute zero. Returns NaN if there are no edges or mean fitness is zero. The sign follows mean fitness.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import gradient_intensity
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> round(gradient_intensity(landscape), 3)
1.143

Neighbor Fitness Correlation

Calculates the correlation between a configuration's fitness and the mean fitness of its neighbors across the fitness landscape.

Neighbor fitness correlation (NFC) reflects how fitness values are distributed spatially. A strong positive correlation indicates that high-fitness configurations tend to be surrounded by other high-fitness neighbors—and similarly for low-fitness ones—suggesting clustering, separability, and an underlying gradient in the landscape. In contrast, a weak or negative correlation implies that high- and low-fitness configurations are intermixed, signaling high fitness variability and a rugged, unstructured landscape.

graphfla.analysis.neighbor_fitness_correlation(
    landscape, auto_calculate=True, method="pearson"
) -> float

Return the correlation between fitness and mean neighbor fitness.

Parameters

landscape : Landscape

Built fitness landscape.

auto_calculate : bool, default=True

Compute missing neighbor-fitness attributes through landscape.neighbor_fitness. If False, require the cached attributes.

method : (pearson, spearman, kendall), default="pearson"

Correlation coefficient to calculate.

Returns

correlation : float

Correlation between fitness and mean neighbor fitness in [-1, 1]. Configurations with missing values are excluded; return NaN if none remain or the correlation is undefined.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import neighbor_fitness_correlation
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> round(neighbor_fitness_correlation(landscape), 3)
-0.169

This metric quantifies the extent to which fitter configurations tend to have neighbors with higher fitness values. A strong positive correlation suggests that higher-fitness configurations exist in higher-fitness regions of the landscape, indicating a structured landscape with potential fitness gradients.

Notes

Positive correlation means fitter configurations tend to have fitter neighbors; negative correlation indicates the opposite association.

Raises

RuntimeError

If auto_calculate=False and neighbor fitness metrics haven't been calculated.

ValueError

If the method is invalid or Pearson correlation has fewer than two valid fitness/neighbor-fitness pairs.