Skip to content

Epistasis

Epistasis – interactions among mutations that produce nonadditive effects on phenotype and fitness – has long been recognized as fundamentally important to understanding the structure and function of genetic pathways and the evolutionary dynamics of complex genetic systems. It plays a major role in evolution by determining the accessibility of mutational pathways and thereby influencing the rate of adaptation and the diversity and robustness of genetic mutants.

In simple terms, it describes the phenomenon that the fitness effect of a mutation depends on what other mutations are already present (i.e., genetic background). GraphFLA provides various analysis tools as introduced in this page to quantify epistasis in empirical data.

Overview

API Purpose
classify_epistasis Proportions of magnitude / sign / reciprocal-sign / positive / negative epistasis via 4-node motifs.
walsh_hadamard Coefficients and contributions by interaction order, using OLS or Lasso.
diminishing_returns_index Correlation between background fitness and improvement size.
increasing_costs_index Correlation between background fitness and detrimental-mutation cost.
gamma Correlation of mutation effects across single-mutant neighbors (Ferretti et al. 2016).
gamma_star Sign-only variant of gamma.
idiosyncratic_index Per-mutation idiosyncrasy (Lyons et al. 2020).
global_idiosyncratic_index Landscape average of idiosyncratic_index across all mutations.
extradimensional_bypass Fraction of reciprocal-sign motifs that admit a fitness-improving bypass.

Basic Types of Epistasis

Calculates the prevalence of five basic types of pairwise epistasis in a fitness landscape.

Description

Specifically, epistatic interactions can be either classified into (e.g., Phillips 2008):

  • Negative epistasis: It occurs when the combined effect of two or more mutations on fitness is less favorable than would be predicted from their individual effects.
    • For deleterious mutations, it means their joint impact results in a synergistic fitness defect that is greater, or more severe, than expected. In extreme cases, this can lead to synthetic lethal or sick interactions, in which a combination of two individually viable mutations causes death (Wang et al. 2017).
    • For beneficial mutations, it means their combined beneficial effect on fitness is smaller than anticipated. This phenomenon is often described as diminishing-returns epistasis (see later) or antagonistic.
  • Positive epistasis: It occurs when the combined effect of two or more mutations on fitness is more favorable than would be predicted from their individual effects.
    • For beneficial mutations, this means their joint impact leads to a synergistic fitness increase that is greater than anticipated.
    • For deleterious mutations, this means their combined detrimental effect on fitness is less severe, or weaker, than expected (i.e., antagonistic). An extreme form of positive interaction is genetic suppression, where the double mutant exhibits better fitness than the least fit single mutant (Leeuwen et al. 2016).

Or (e.g., see Poelwijk 2007):

  • Magnitude epistasis: In magnitude epistasis, the effect of a mutation (e.g., beneficial or deleterious) maintains its sign (direction) regardless of the genetic background. However, the magnitude (strength) of this effect changes depending on what other mutations are present. For example, a mutation might be strongly beneficial in one genetic background but only weakly beneficial in another, yet it remains beneficial in both contexts. The fitness effects are non-additive, but the qualitative outcome (beneficial/deleterious) for a given mutation does not change.
  • Sign epistasis: Sign epistasis occurs when the effect of a mutation changes its sign (i.e., from beneficial to deleterious, or vice versa) depending on the genetic background (the presence of other mutations). A mutation that is advantageous in one genetic context might be detrimental in another, or a deleterious mutation might become beneficial when another specific mutation is present. This type of epistasis is significant because it can create rugged fitness landscapes with multiple fitness peaks, potentially trapping evolving populations at suboptimal states as some evolutionary paths become inaccessible.
  • Reciprocal sign epistasis: Reciprocal sign epistasis is a specific and stronger form of sign epistasis. It occurs when the sign of the fitness effect of each of two interacting mutations is reversed by the presence of the other mutation. A common scenario is where two mutations are individually deleterious (or less fit than the wild type), but their combination results in a genotype that is more fit than either single mutant, and potentially even more fit than the original wild type. This type of interaction is a necessary condition for the existence of multiple peaks on a fitness landscape.

This implementation (from Papkou et al., 2023) calculates the fraction of each of these 5 types of epistasis among all pairs of mutations (i.e., pairwise epistasis) using the motif function in igraph.

Note

The two sets of fractions summarize the directed motifs found in the constructed graph. Their normalization and behavior when no relevant motifs are found are described in Returns below.

Warning

This method can be quite computationally expensive for large landscapes (e.g., >10,000 mutants). Use sample_cut_prob="auto" for automatic sampling, or choose a sampling probability explicitly. Set seed for reproducible sampling.

graphfla.analysis.classify_epistasis(
    landscape,
    sample_cut_prob="auto",
    seed=None,
    time_budget=15.0,
) -> Dict[str, float]

Return proportions of five epistasis types among directed four-node motifs.

Parameters

landscape : Landscape

The fitness landscape object, containing landscape.graph as an igraph.Graph with a "fitness" vertex attribute.

sample_cut_prob : auto, default="auto"

Controls the 4-motif search. "auto" (default) picks a pruning probability from a runtime ladder targeting time_budget (small landscapes resolve to exact). This is not a hard time limit. 0 or None forces exact enumeration; a float in (0, 1] sets the igraph pruning probability directly (higher -> faster, less accurate).

seed : int or None, default=None

Seed for the sampling RNG, making approximate results reproducible.

time_budget : float, default=15.0

Target wall-clock seconds for sample_cut_prob="auto".

Returns

result : dict of str to float

Proportions keyed by epistasis type:

  • magnitude: The magnitude of the combined fitness effect of mutations differs from the sum of their individual effects, but the direction relative to single mutants or wild-type may not change sign.

  • sign: The sign of the fitness effect of at least one mutation changes depending on the presence of other mutations. For example, a mutation beneficial on its own becomes deleterious when combined with another specific mutation.

  • reciprocal_sign: A specific form of sign epistasis where the sign of the effect of each mutation depends on the allele state at the other locus.

  • positive: The combined fitness effect of mutations is greater than the sum of their individual effects, often referred to as synergistic epistasis.

  • negative: The combined fitness effect of mutations is less than the sum of their individual effects, often referred to as antagonistic epistasis.

Values are zero if relevant counts/instances are zero or cannot be processed.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import classify_epistasis
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> classify_epistasis(landscape, sample_cut_prob=0)["magnitude"]
1.0

Determines magnitude, sign, and reciprocal sign epistasis based on counts/estimates of motifs 19, 52, 66. Determines positive and negative epistasis by analyzing the fitness relationships within instances of these motifs.

Notes

Fractions are normalized over the directed motifs found in the constructed graph. Missing or neutral edges can remove a square from this population. The relation gamma_star = 1 - sign - 2*reciprocal_sign requires exact counts of the same variable squares with no neutral effects; it is not an identity for arbitrarily filtered or sampled graphs. See gamma_star.

Raises

AttributeError

If landscape.graph is not an igraph.Graph object or does not exist.

ValueError

If sample_cut_prob is invalid, or if the fitness attribute is missing.

Higher-order Epistasis

In addition to pairwise epistasis, higher-order epistasis that involves multiple mutations are also common in empirical fitness landscapes (e.g., Weinreich et al. 2013, Domingo et al. 2018). GraphFLA uses walsh_hadamard to estimate coefficients and summarize contributions by interaction order.

Walsh-Hadamard coefficients

To identify which variables interact, fitness can be represented as a sum of individual effects and interactions. For binary variables, let \(z_i=x_i-\tfrac12\), where \(x_i\in\{0,1\}\). The fitted landscape takes the form

\[ \widehat f(\mathbf{x})=\varepsilon_0 +\sum_i\varepsilon_i z_i +\sum_{i<j}\varepsilon_{ij}z_i z_j+\cdots. \]

Here, \(\varepsilon_0\) is the model's mean over uniformly weighted configurations. First-order coefficients describe background-averaged effects of individual changes, second-order coefficients describe pairwise interactions, and higher orders capture effects involving additional variables. The centered encoding determines the scaling of these coefficients.

GraphFLA implements the multistate extension of Faure et al. (2024). For a variable with \(s_i\) states, \(z_i\) is replaced by centered indicators \(\phi_{i,a}(x_i)=\mathbf{1}[x_i=a]-1/s_i\), one for each nonreference state \(a\). Interaction terms multiply indicators from distinct variables.

One call returns coefficients and an order summary. Cumulative R² and its increments describe refitted models through each order; the model variance fractions describe the highest-order fit over uniformly weighted backgrounds. Lasso provides regularized estimates. Its model spectrum and observed-data fit gains answer different questions and are reported separately.

graphfla.analysis.walsh_hadamard(
    landscape,
    max_order=2,
    max_cells=10000000.0,
    chunk_size=1000,
    *,
    method="ols",
    alpha="cv",
    cv=5,
    random_state=0,
    max_iter=10000,
    tol=0.0001,
    n_jobs=1,
) -> Dict[str, Union[pd.DataFrame, dict]]

Return fitted Walsh-Hadamard coefficients and contributions by order.

Parameters

landscape : Landscape

Built landscape with finite fitness and unique, nonmissing configurations. Use the nodes retained in get_data(); construction filters are not reversed. All variable types, including ordinal variables, are treated as discrete states without numeric spacing. States absent from the retained data are not inferred. Fitness is used in its supplied units, without changing sign for minimization.

max_order : int, default=2

Largest number of distinct interacting variables. Must be nonnegative. Zero fits only a constant; values above the number of variable sites include all orders. Truncation specifies a model, not evidence that omitted interactions are absent.

max_cells : int or float, default=1e7

Maximum number of entries in the dense design, including its constant column. Checked before enumerating terms or allocating the design. Each entry uses eight bytes. This is not a total memory limit: solver copies and workspace also require memory.

chunk_size : int, default=1000

Maximum number of rows in temporary feature-construction blocks. The final design is dense regardless of this value.

method : (ols, lasso), default="ols"

Coefficient estimator. OLS raises on a rank-deficient design. Lasso minimizes squared prediction error plus an L1 coefficient penalty; it does not establish identifiability of unregularized coefficients. Neither method automatically switches to the other.

alpha : float or {cv}, default="cv"

Positive Lasso penalty in the objective sum((y - prediction)**2) / (2 * n_samples) + alpha * sum(abs(coef)). The constant term is not penalized and features are not standardized. "cv" selects among 100 logarithmically spaced penalties from the smallest all-zero-slope penalty to one thousandth of that value, using mean validation squared error. Ignored for OLS.

cv : int, default=5

Number of shuffled folds when method="lasso", alpha="cv". Must be between 2 and the number of retained observations. The selected model is refitted on all observations; no held-out score is returned.

random_state : (int, numpy.random.RandomState or None), default=0

Controls shuffled cross-validation folds. Pass an integer for reproducible splits. Used only for cross-validated Lasso.

max_iter : int, default=10000

Maximum coordinate-descent iterations for Lasso.

tol : float, default=1e-4

Positive convergence tolerance for Lasso.

n_jobs : int or None, default=1

Number of parallel cross-validation jobs for Lasso. -1 uses all available processors. Does not parallelize the sequence of orders.

Returns

result : dict

Dictionary with coefficients, order_summary and fit_info. result["coefficients"] is a DataFrame sorted by (order, term) with:

  • order: number of interacting variables.

  • positions: tuple of one-based original feature positions.

  • term: label source_position_target, joined by - for interactions. Allele-label characters %, _ and - are percent-escaped; ordinary sequence labels are unchanged.

  • coefficient: fitted effect in fitness units.

The order-zero row is labeled intercept. It is the model's uniform full-space mean, not the reference configuration's fitness. result["order_summary"] has one row per order from zero through the effective maximum, with columns:

  • order: maximum order in the corresponding nested fit.

  • r2: cumulative training R-squared of that fit.

  • delta_r2: increase over the preceding order (zero at order zero).

  • rmse: training root mean squared error in fitness units.

  • n_terms and rank: design width and OLS rank (NA for Lasso).

  • alpha: penalty selected for that nested fit (NaN for OLS or the constant baseline; zero for a constant CV solution).

  • model_variance_fraction: that order's share of the highest-order fitted model's variance over uniform product-space backgrounds. This is a model spectrum, not an observed-data R-squared increment.

  • n_nonzero: number of nonzero coefficients of exactly that order in the highest-order Lasso fit (NA for OLS).

R-squared and its increments are NaN for constant fitness. Model variance fractions are NaN for a constant fitted model. Negative regularized R-squared increments are retained, not clipped. result["fit_info"] records the estimator, dimensions, final rank and alpha, reference alleles, original column names and scoring conventions. The same metadata is attached to both tables; save it separately when exporting to formats such as CSV.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import walsh_hadamard
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [10., 13., 12., 20.], verbose=False
... )
>>> result = walsh_hadamard(landscape)
>>> coefficients = result["coefficients"]
>>> round(float(coefficients.set_index("term").loc["0_1_1-0_2_1", "coefficient"]), 6)
5.0
>>> regularized = walsh_hadamard(landscape, method="lasso", alpha=0.1)
>>> regularized["fit_info"]["alpha"]
0.1
>>> result["order_summary"].order.tolist()
[0, 1, 2]

Use the multistate Walsh-Hadamard representation of Faure et al. 1. Fit terms through max_order by ordinary least squares or explicit Lasso regularization. Incomplete landscapes are allowed; OLS requires a design matrix with full column rank. Build the design once and reuse it to fit the nested models in the order summary.

Notes

Use observed allele 0 as reference for Boolean variables and the first retained observation for other variables. Construct columns of \(H^{-1}V^{-1}\) as products of centered state indicators, then fit equally weighted observations 1. Backgrounds are uniform over the Cartesian product of observed states. On a complete landscape with all orders, OLS equals the direct transform; incomplete or truncated fits estimate model coefficients.

Refit each nested model; CV selects alpha within each order using the same folds. Scores describe training fit, not held-out prediction or causal attribution. The spectrum uses exact product-space covariance, accounting for correlated multistate columns. Lasso leaves the constant unpenalized; this and the adaptive CV grid differ from the authors' analysis scripts.

References

[1] Faure, A. J., Lehner, B., Miro Pina, V., Serrano Colome, C., and Weghorn, D. (2024). An extension of the Walsh-Hadamard transform to calculate and model epistasis in genetic landscapes of arbitrary shape and complexity. PLOS Computational Biology 20(5): e1012132. https://doi.org/10.1371/journal.pcbi.1012132. Eqs. (9), (12), (23)-(25).

Raises

graphfla.exceptions.NotBuiltError

If the landscape has not been built.

ValueError

If inputs or parameters are invalid, the design exceeds max_cells, OLS coefficients are not identifiable, or coefficients exceed float64.

numpy.linalg.LinAlgError

If the least-squares decomposition fails to converge.

Warns

UserWarning

If fitness is constant and R-squared is undefined.

sklearn.exceptions.ConvergenceWarning

If Lasso does not converge within max_iter iterations.

Global Epistasis

Beyond interactions between specific pairs or sets of mutations, global epistasis describes a broader pattern where the fitness effect of a mutation systematically depends on the overall fitness of the genetic background in which it arises. This often approximately one-dimensional relationship can emerge from the combined effects of many specific, idiosyncratic interactions, which may also explain variation around the overall trend.

Common forms of these fitness-correlated trends are diminishing returns for beneficial mutations and increasing costs for deleterious mutations, with the same mutation having a less positive or more negative effect on fitter backgrounds. The returns and costs indices pool different mutations, so their signs alone do not establish mutation-specific epistasis or statistical significance.

Diminishing Returns Epistasis

Diminishing returns epistasis is a form of global epistasis that applies to beneficial mutations. It describes the pattern where the positive fitness effect of a mutation is smaller when it occurs in an already fit genetic background than in a less fit background. As beneficial mutations accumulate, shrinking gains can contribute to decelerating adaptation (Chou et al., 2011).

GraphFLA calculates the correlation or regression slope between background fitness and the positive fitness increase of beneficial transitions, with each retained transition contributing equally. A negative value describes diminishing gains across this pooled set. The result depends on the fitness scale and the mutations available in each background.

graphfla.analysis.diminishing_returns_index(
    landscape,
    method: Literal[
        "pearson", "spearman", "regression"
    ] = "pearson",
) -> float

Return the pooled trend of beneficial effects with background fitness.

Parameters

landscape : Landscape

Built landscape with a directed graph of improving transitions and finite node fitness. Use a one-step neighborhood for mutation-level interpretation. The graph's retained edges determine the population; construction filters and missing configurations are not undone.

method : (pearson, spearman, regression), default="pearson"

Pearson correlation, Spearman correlation with average ranks for ties, or the ordinary least-squares slope with an intercept. Each edge has equal weight. Effects are differences on the supplied fitness scale.

Returns

index : float

Trend between starting fitness and the positive improvement for each retained edge. Negative values describe smaller gains on better backgrounds. Fitness is negated for minimization, so the interpretation is unchanged. Returns NaN for fewer than two eligible edges or constant starting fitness. A constant effect gives NaN for correlation and zero for regression. No significance test or p-value is returned.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import diminishing_returns_index
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0., 2., 3., 4.],
...     epsilon=0, verbose=False
... )
>>> round(diminishing_returns_index(landscape, method="regression"), 3)
-0.444

Notes

For each stored improving edge u -> v, correlate q(u) with q(v) - q(u), where q = fitness for maximization and q = -fitness for minimization. Zero-effect edges are excluded. This pools individual transitions 1, rather than averaging effects per node or tracking a fixed mutation across backgrounds. Edge attributes such as delta_fit are not used.

The result is descriptive: effect-sign selection, differing mutation composition (even in an additive landscape) and shared measurement error can produce a trend without demonstrating mutation-specific global epistasis. Fitness transformations can change it. Pearson and regression take O(V + E) time and O(V + B) auxiliary memory with fixed block size B; exact Spearman requires O(E) additional memory and O(E log E) time.

References

[1] Huang, M., Zhou, S. and Li, K. (2025). Augmenting Biological Fitness Prediction Benchmarks with Landscapes Features from GraphFLA. NeurIPS 38, Appendix C.3.2. https://doi.org/10.52202/085713-1180.

[2] Papkou, A. et al. (2023). A rugged yet easily navigable fitness landscape. Science 382, eadh3860, Fig. S22. https://doi.org/10.1126/science.adh3860.

Raises

RuntimeError

If the landscape has not been built.

ValueError

If method is invalid, the graph or fitness is missing, fitness is nonfinite, or the graph is undirected or contains worsening edges.

Warns

UserWarning

If the requested statistic is undefined.

Increasing Cost Epistasis

Increasing cost epistasis is the analog of diminishing returns for deleterious mutations. It describes a pattern where a mutation's fitness cost becomes more severe in fitter backgrounds. The same mutation may have a weaker detrimental effect in less fit backgrounds (Johnson et al., 2019).

GraphFLA correlates background fitness with the positive magnitude of fitness loss, or fits a regression slope, weighting each retained deleterious transition equally. A positive value describes increasing costs. Each transition reverses a stored improving edge and starts at its better endpoint.

graphfla.analysis.increasing_costs_index(
    landscape,
    method: Literal[
        "pearson", "spearman", "regression"
    ] = "pearson",
) -> float

Return the pooled trend of deleterious cost with background fitness.

Parameters

landscape : Landscape

Built landscape with a directed graph of improving transitions and finite node fitness. Reverse each retained edge to represent a worsening move. A mutation-level interpretation requires a reversible, one-step neighborhood. Construction filters and missing configurations remain part of the input population.

method : (pearson, spearman, regression), default="pearson"

Pearson correlation, Spearman correlation with average ranks for ties, or the ordinary least-squares slope with an intercept. Each reverse transition has equal weight; its cost is a positive fitness difference.

Returns

index : float

Trend between the better endpoint's fitness and the cost of moving to its worse neighbor. Positive values describe larger costs on better backgrounds. Fitness is negated for minimization. Returns NaN for fewer than two eligible edges or constant background fitness. A constant cost gives NaN for correlation and zero for regression. No significance test or p-value is returned.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import increasing_costs_index
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0., 2., 3., 4.],
...     epsilon=0, verbose=False
... )
>>> round(increasing_costs_index(landscape, method="regression"), 3)
-0.364

Notes

For each stored improving edge u -> v, correlate q(v) with q(v) - q(u), using the same oriented fitness and edge population as diminishing_returns_index. Zero-effect edges are excluded. This is the pooled cost convention of 1; it is not the distribution of mutation-specific regressions studied by Johnson et al. 2. Its sign alone does not establish global epistasis or statistical significance.

References

[1] Huang, M., Zhou, S. and Li, K. (2025). Augmenting Biological Fitness Prediction Benchmarks with Landscapes Features from GraphFLA. NeurIPS 38, Appendix C.3.2. https://doi.org/10.52202/085713-1180.

[2] Johnson, M. S. et al. (2019). Higher-fitness yeast genotypes are less robust to deleterious mutations. Science 366, 490-493, Figs. 3-4. https://doi.org/10.1126/science.aay4199.

Raises

RuntimeError

If the landscape has not been built.

ValueError

If method is invalid, the graph or fitness is missing, fitness is nonfinite, or the graph is undirected or contains worsening edges.

Warns

UserWarning

If the requested statistic is undefined.

Gamma Statistic

Gamma

The gamma statistic (\(\gamma\)) measures the amount of epistasis in a fitness landscape through the correlation between fitness effects of mutations across different genetic backgrounds (Ferretti et al., 2016). It compares the effects of the same mutation in backgrounds differing by a single mutation at another locus. In simpler terms, it quantifies how the effect of a particular mutation is altered by another mutation in the genetic background, averaged across the landscape. For individual two-locus, two-allele squares, the following relationships hold; the ranges overlap and do not define classification thresholds for an entire landscape:

  • \(\gamma = 1\): no epistasis; mutation effects are unchanged across backgrounds.
  • Magnitude epistasis \(\rightarrow 0 \leq \gamma < 1\): mutation effects retain their signs but differ in magnitude, giving a nonnegative correlation.
  • Sign epistasis \(\rightarrow -\frac{1}{3} \leq \gamma < 1\): the effect of one mutation changes sign, and the correlation can be positive, zero, or negative.
  • Reciprocal sign epistasis \(\rightarrow -1 \leq \gamma < 0\): the effects of both mutations change sign, giving a negative correlation.
graphfla.analysis.gamma(landscape, n_jobs=-1) -> float

Return the correlation of mutation effects across neighboring backgrounds.

Parameters

landscape : Landscape

Built landscape. Fitness is used on its supplied scale; apply any scientifically appropriate log transformation before construction. All observed allele pairs at two distinct variables are considered. Retained configurations define the population, independently of graph edges, ordinal step restrictions and construction epsilon.

n_jobs : int or None, default=-1

Number of joblib workers. -1 uses all available CPUs and 1 runs serially. None follows the active joblib configuration, or uses one worker without a configuration. Zero is invalid.

Returns

gamma_value : float

Non-centered correlation in [-1, 1]. One means equal mutation effects across every observed square; negative values indicate opposing effects. Returns NaN if there are no complete squares or every effect on those squares is zero.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import gamma
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> round(gamma(landscape, n_jobs=1), 6)
0.888889

Notes

For parallel effects a and b, pool their products and squared effects: sum(a*b) / sum((a*a + b*b)/2) over both directions of every complete two-variable square 1. Each allele-pair combination has equal weight; averaging square ratios gives a different statistic. Fitness offsets and nonzero linear rescaling leave gamma unchanged.

Missing corners exclude a square from both sums. This describes the observed squares, not the paper's separate distance-correlation estimator for missing data (Eq. (2)). A value of one on incomplete data does not establish global additivity.

References

[1] Ferretti, L. et al. (2016). Measuring epistasis in fitness landscapes: The correlation of fitness effects of mutations. Journal of Theoretical Biology, 396, 132-143. Eqs. (1), (3), Appendix C.1. https://doi.org/10.1016/j.jtbi.2016.01.037

Raises

graphfla.exceptions.NotBuiltError

If the landscape has not been built.

ValueError

If the graph lacks fitness values, or the worker count is invalid.

Warns

UserWarning

If fewer than two variables remain in the built landscape.

Gamma Star

The gamma-star statistic (\(\gamma^*\)) is a variant of \(\gamma\) that focuses only on sign consistency, ignoring the magnitude of fitness effects. It indicates whether mutations tend to have consistent directional effects across different genetic backgrounds.

graphfla.analysis.gamma_star(landscape, n_jobs=-1) -> float

Return the correlation of mutation-effect signs across backgrounds.

Parameters

landscape : Landscape

Built landscape. Positive, zero and negative fitness differences are assigned +1, 0 and -1, respectively. Only exact ties are neutral; construction epsilon does not set a sign tolerance for this metric.

n_jobs : int or None, default=-1

Number of joblib workers. -1 uses all available CPUs and 1 runs serially. None follows the active joblib configuration, or uses one worker without a configuration. Zero is invalid.

Returns

gamma_star_value : float

Sign correlation in [-1, 1]. One indicates consistent nonzero signs, minus one indicates reversed signs, and zero indicates cancellation or absence of nonzero parallel products. Returns NaN if there are no complete squares or all effects on those squares are neutral.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import gamma_star
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 0, 1, 2], verbose=False)
>>> round(gamma_star(landscape, n_jobs=1), 6)
0.666667

Notes

Apply the pooling formula of gamma to effect signs 1. A zero effect contributes zero to the numerator and its squared-effect term; it does not cause the entire square to be discarded. The function uses Eq. (11) with zero tolerance, not the optional positive-tolerance variant.

The identity gamma_star = 1 - phi_sign - 2*phi_reciprocal (Eq. (12)) requires no neutral effects and the same square population. It need not hold for classify_epistasis when graph filtering removes edges, graph motifs differ from variable squares, or motif counts are sampled.

References

[1] Ferretti, L. et al. (2016). Measuring epistasis in fitness landscapes: The correlation of fitness effects of mutations. Journal of Theoretical Biology, 396, 132-143. Eqs. (10)-(12), Appendix C.3. https://doi.org/10.1016/j.jtbi.2016.01.037

Raises

graphfla.exceptions.NotBuiltError

If the landscape has not been built.

ValueError

If the graph lacks fitness values, or the worker count is invalid.

Warns

UserWarning

If fewer than two variables remain in the built landscape.

Idiosyncratic Epistasis

The fitness-correlated trends described above summarize how mutation effects change with background fitness. Individual mutations can also depend on the specific alleles present in a background. For example, a substitution that improves a protein's activity in one sequence may become harmful after another residue changes. Such interactions are described as idiosyncratic epistasis and help explain why a mutation's effect in one genotype may not predict its effect in another.

Lyons et al. (2020) introduced the idiosyncratic index to quantify background dependence. For a given mutation, it compares the standard deviation of its effects across matching backgrounds with that of fitness differences in an equally sized sample of random genotype pairs. GraphFLA summarizes the landscape by averaging these ratios across eligible directed mutations, giving each mutation equal weight.

When defined, an index of zero indicates constant mutation effects across the observed backgrounds. A value near one indicates variation comparable to random genotype-pair differences; estimates can exceed one. The index measures variation, rather than the fraction of epistatic mutations. Because a nonlinear global fitness relationship can also generate background dependence, a positive index alone does not distinguish specific interactions from global epistasis.

Per-mutation Index

graphfla.analysis.idiosyncratic_index(
    landscape, mutation, min_pairs: int = 3, *, seed=None
) -> float

Return a mutation's idiosyncratic index from matched backgrounds.

Parameters

landscape : Landscape

Built landscape with unique configurations and finite fitness values. Both matched backgrounds and random controls use the genotypes retained in landscape.get_data().

mutation : tuple of (source, position, target)

Allele substitution to evaluate. position is a configuration-column label from landscape.data_types, not a positional column index. source and target must be distinct observed alleles at that position. For example, (0, "bit_0", 1) changes the first Boolean feature from 0 to 1.

min_pairs : int, default=3

Minimum number of observed matching backgrounds, at least 2. All matching backgrounds are used when this threshold is met. The control then contains the same number of random pairs. This threshold is a GraphFLA estimation guard, not a cutoff specified in 1.

seed : int or None, default=None

Seed for a local NumPy RandomState. An integer reproduces the control sample for the same ordered input; None starts a fresh random stream. NumPy's global RNG state is not modified.

Returns

index : float

Ratio of observed-effect SD to sampled-control SD. Values can exceed 1. Returns NaN for a constant-fitness landscape, fewer than min_pairs backgrounds, or a sampled control with zero SD. A well-defined value of zero indicates constant mutation effects across the observed backgrounds.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import idiosyncratic_index
>>> sequences = ["000", "001", "010", "011", "100", "101", "110", "111"]
>>> fitness = [0., 1., 2., 3., 1., 2., 4., 6.]
>>> landscape = BooleanLandscape().build_from_data(
...     sequences, fitness, epsilon=0, verbose=False
... )
>>> value = idiosyncratic_index(landscape, (0, "bit_0", 1), seed=0)
>>> round(value, 3)
0.308

The index compares the standard deviation of one mutation's effects with that of an equally sized sample of random genotype-pair differences 1. Effects and controls use the fitness scale supplied in the landscape.

Notes

Match genotypes at all positions except the focal one. Divide the SD of their fitness differences by that of equally many random-pair differences, sampling both endpoints independently with replacement. Both SDs use ddof=0; matching does not depend on graph edges.

With seed=None, results can differ across calls. Fitness is used as supplied. The index measures background dependence, including variation that can arise from a nonlinear global fitness map.

References

[1] Lyons, Daniel M., Zhengting Zou, Haiqing Xu, and Jianzhi Zhang. "Idiosyncratic epistasis creates universals in mutational effects and evolutionary trajectories." Nature Ecology & Evolution 4 (2020): 1685-1693. https://doi.org/10.1038/s41559-020-01286-y.

Raises

graphfla.exceptions.NotBuiltError

If the landscape has not been built.

ValueError

If min_pairs or seed is invalid, the mutation uses an unknown position or allele, or source and target are equal. Also raised for missing configuration values, duplicate configurations, or nonfinite fitness.

Warns

RuntimeWarning

If the sampled control has zero SD. No replacement sample is drawn.

Global (Landscape-Wide) Index

graphfla.analysis.global_idiosyncratic_index(
    landscape, n_jobs=-1, seed=None, min_pairs: int = 3
) -> float

Return the mean idiosyncratic index across directed mutations.

Parameters

landscape : Landscape

Built landscape with unique configurations and finite fitness values. Mutation effects and random controls both use the genotypes retained in landscape.get_data(). Apply the intended fitness transformation and population selection before constructing the landscape.

n_jobs : int or None, default=-1

Number of parallel jobs for background matching. -1 uses all available cores; 1 runs serially. This does not change a fixed-seed result.

seed : int or None, default=None

Seed for a local NumPy RandomState. An integer reproduces the result for the same ordered input and parameters. None starts a fresh random stream. NumPy's global RNG state is not modified.

min_pairs : int, default=3

Minimum number of matching backgrounds for a mutation to contribute, at least 2. Mutations below this threshold are omitted. All matching backgrounds of eligible mutations are used, with equally sized random controls. The paper specifies no minimum for this index.

Returns

mean_index : float

Arithmetic mean of the eligible mutation indices. Values can exceed 1. Returns NaN for constant fitness, no eligible mutations, or a zero-SD control for any eligible mutation. A failed control is not omitted from the mean or resampled.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import global_idiosyncratic_index
>>> sequences = ["000", "001", "010", "011", "100", "101", "110", "111"]
>>> fitness = [0., 1., 2., 3., 1., 2., 4., 6.]
>>> landscape = BooleanLandscape().build_from_data(
...     sequences, fitness, epsilon=0, verbose=False
... )
>>> round(global_idiosyncratic_index(landscape, n_jobs=1, seed=0), 3)
0.416

Apply the SD ratio defined by Lyons et al. 1 to each eligible directed mutation and return their arithmetic mean. Every mutation receives equal weight, irrespective of its number of observed backgrounds.

Notes

Apply idiosyncratic_index's calculation to each ordered allele pair at each position. Forward and reverse mutations receive independent controls. Average over eligible mutations, not positions or backgrounds.

Matching uses configuration values rather than graph edges. Controls use only retained genotypes, so graph pruning changes the reference population.

References

[1] Lyons, Daniel M., Zhengting Zou, Haiqing Xu, and Jianzhi Zhang. "Idiosyncratic epistasis creates universals in mutational effects and evolutionary trajectories." Nature Ecology & Evolution 4 (2020): 1685-1693. https://doi.org/10.1038/s41559-020-01286-y.

Raises

graphfla.exceptions.NotBuiltError

If the landscape has not been built.

ValueError

If min_pairs or the random seed is invalid, or configurations are missing or duplicated, or fitness contains nonfinite values.

Warns

RuntimeWarning

If any eligible mutation's sampled control has zero SD.

Extradimensional Bypass Analysis

Detects extradimensional bypasses around reciprocal-sign-epistasis (RSE) motifs.

Description

Reciprocal sign epistasis occurs when the wildtype (ab) and the double mutant (AB) are both fitter than the two intermediate single mutants (aB, Ab). This creates a fitness valley along the direct path from ab to AB that natural selection cannot cross. However, when the landscape has more than two relevant dimensions, the population can sometimes traverse a third-position intermediate (i.e., gain and later lose an extra mutation) to bypass the valley. Such indirect uphill paths are called extradimensional bypasses and they greatly affect a landscape's effective navigability.

For each RSE motif in the landscape (igraph isomorphism class 19), this function checks whether the directed improving graph admits any path from ab to AB. The fraction of RSE motifs that do admit such a bypass is the bypass proportion.

graphfla.analysis.extradimensional_bypass(
    landscape,
    sample_cut_prob="auto",
    seed=None,
    time_budget=15.0,
) -> Dict[str, Union[float, int]]

Return a summary of bypasses around reciprocal-sign epistasis motifs.

Parameters

landscape : Landscape

The fitness landscape object, containing landscape.graph as an igraph.Graph with a "fitness" vertex attribute.

sample_cut_prob : auto, default="auto"

Controls the 4-motif search. "auto" (default) picks a pruning probability from a runtime ladder targeting time_budget (small landscapes resolve to exact). This is not a hard time limit. 0 or None forces exact enumeration; a float in (0, 1] sets the igraph pruning probability directly (higher -> faster, less accurate).

seed : int or None, default=None

Seed for the sampling RNG, making approximate results reproducible.

time_budget : float, default=15.0

Target wall-clock seconds for sample_cut_prob="auto".

Returns

result : dict

Summary with the following keys:

  • bypass_proportion : proportion of reciprocal-sign-epistasis motifs for which an extradimensional bypass exists (float between 0 and 1).

  • average_bypass_length : average length of extradimensional bypasses for motifs where such bypasses exist; NaN if none exist.

  • total_motifs : integer total number of type-19 motifs analyzed.

  • motifs_with_bypass : integer number of motifs that have a bypass.

Examples

>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import extradimensional_bypass
>>> landscape = BooleanLandscape().build_from_data(
...     ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> extradimensional_bypass(landscape, sample_cut_prob=0)["bypass_proportion"]
0.0

For each motif representing reciprocal sign epistasis (type 19), this function identifies whether accessible evolutionary paths exist that bypass the direct path between the double mutant nodes. Such indirect paths are called extradimensional bypasses and allow evolution to traverse fitness valleys that would otherwise be inaccessible under strong selection.

Notes

Reciprocal sign epistasis occurs when both the wildtype (ab) and double mutant (AB) have higher fitness than both single mutants (aB, Ab). This creates a fitness valley that prevents direct evolutionary access between ab and AB. Extradimensional bypasses are indirect paths through the broader fitness landscape that circumvent this valley.

Raises

AttributeError

If landscape.graph is not an igraph.Graph object or does not exist.

ValueError

If sample_cut_prob is not between 0 and 1, or if fitness attribute missing.