Epistasis
Epistasis – interactions among mutations that produce nonadditive effects on phenotype and fitness – has long been recognized as fundamentally important to understanding the structure and function of genetic pathways and the evolutionary dynamics of complex genetic systems. It plays a major role in evolution by determining the accessibility of mutational pathways and thereby influencing the rate of adaptation and the diversity and robustness of genetic mutants.
In simple terms, it describes the phenomenon that the fitness effect of a mutation depends on what other mutations are already present (i.e., genetic background). GraphFLA provides various analysis tools as introduced in this page to quantify epistasis in empirical data.
Overview¶
| API | Purpose |
|---|---|
classify_epistasis |
Proportions of magnitude / sign / reciprocal-sign / positive / negative epistasis via 4-node motifs. |
walsh_hadamard |
Coefficients and contributions by interaction order, using OLS or Lasso. |
diminishing_returns_index |
Correlation between background fitness and improvement size. |
increasing_costs_index |
Correlation between background fitness and detrimental-mutation cost. |
gamma |
Correlation of mutation effects across single-mutant neighbors (Ferretti et al. 2016). |
gamma_star |
Sign-only variant of gamma. |
idiosyncratic_index |
Per-mutation idiosyncrasy (Lyons et al. 2020). |
global_idiosyncratic_index |
Landscape average of idiosyncratic_index across all mutations. |
extradimensional_bypass |
Fraction of reciprocal-sign motifs that admit a fitness-improving bypass. |
Basic Types of Epistasis¶
Calculates the prevalence of five basic types of pairwise epistasis in a fitness landscape.
Description
Specifically, epistatic interactions can be either classified into (e.g., Phillips 2008):
- Negative epistasis: It occurs when the combined effect of two or more mutations on fitness is less favorable than would be predicted from their individual effects.
- For deleterious mutations, it means their joint impact results in a synergistic fitness defect that is greater, or more severe, than expected. In extreme cases, this can lead to synthetic lethal or sick interactions, in which a combination of two individually viable mutations causes death (Wang et al. 2017).
- For beneficial mutations, it means their combined beneficial effect on fitness is smaller than anticipated. This phenomenon is often described as diminishing-returns epistasis (see later) or antagonistic.
- Positive epistasis: It occurs when the combined effect of two or more mutations on fitness is more favorable than would be predicted from their individual effects.
- For beneficial mutations, this means their joint impact leads to a synergistic fitness increase that is greater than anticipated.
- For deleterious mutations, this means their combined detrimental effect on fitness is less severe, or weaker, than expected (i.e., antagonistic). An extreme form of positive interaction is genetic suppression, where the double mutant exhibits better fitness than the least fit single mutant (Leeuwen et al. 2016).
Or (e.g., see Poelwijk 2007):
- Magnitude epistasis: In magnitude epistasis, the effect of a mutation (e.g., beneficial or deleterious) maintains its sign (direction) regardless of the genetic background. However, the magnitude (strength) of this effect changes depending on what other mutations are present. For example, a mutation might be strongly beneficial in one genetic background but only weakly beneficial in another, yet it remains beneficial in both contexts. The fitness effects are non-additive, but the qualitative outcome (beneficial/deleterious) for a given mutation does not change.
- Sign epistasis: Sign epistasis occurs when the effect of a mutation changes its sign (i.e., from beneficial to deleterious, or vice versa) depending on the genetic background (the presence of other mutations). A mutation that is advantageous in one genetic context might be detrimental in another, or a deleterious mutation might become beneficial when another specific mutation is present. This type of epistasis is significant because it can create rugged fitness landscapes with multiple fitness peaks, potentially trapping evolving populations at suboptimal states as some evolutionary paths become inaccessible.
- Reciprocal sign epistasis: Reciprocal sign epistasis is a specific and stronger form of sign epistasis. It occurs when the sign of the fitness effect of each of two interacting mutations is reversed by the presence of the other mutation. A common scenario is where two mutations are individually deleterious (or less fit than the wild type), but their combination results in a genotype that is more fit than either single mutant, and potentially even more fit than the original wild type. This type of interaction is a necessary condition for the existence of multiple peaks on a fitness landscape.
This implementation (from Papkou et al., 2023) calculates the fraction of each of these 5 types of epistasis among all pairs of mutations (i.e., pairwise epistasis) using the motif function in igraph.
Note
The two sets of fractions summarize the directed motifs found in the constructed graph. Their normalization and behavior when no relevant motifs are found are described in Returns below.
Warning
This method can be quite computationally expensive for large landscapes (e.g., >10,000 mutants). Use sample_cut_prob="auto" for automatic sampling, or choose a sampling probability explicitly. Set seed for reproducible sampling.
graphfla.analysis.classify_epistasis(
landscape,
sample_cut_prob="auto",
seed=None,
time_budget=15.0,
) -> Dict[str, float]Return proportions of five epistasis types among directed four-node motifs.
Parameters
-
landscape: Landscape The fitness landscape object, containing landscape.graph as an igraph.Graph with a "fitness" vertex attribute.
-
sample_cut_prob: auto, default="auto" Controls the 4-motif search.
"auto"(default) picks a pruning probability from a runtime ladder targetingtime_budget(small landscapes resolve to exact). This is not a hard time limit.0orNoneforces exact enumeration; a float in(0, 1]sets the igraph pruning probability directly (higher -> faster, less accurate).-
seed: int or None, default=None Seed for the sampling RNG, making approximate results reproducible.
-
time_budget: float, default=15.0 Target wall-clock seconds for
sample_cut_prob="auto".
Returns
-
result: dict of str to float Proportions keyed by epistasis type:
-
magnitude: The magnitude of the combined fitness effect of mutations differs from the sum of their individual effects, but the direction relative to single mutants or wild-type may not change sign. -
sign: The sign of the fitness effect of at least one mutation changes depending on the presence of other mutations. For example, a mutation beneficial on its own becomes deleterious when combined with another specific mutation. -
reciprocal_sign: A specific form of sign epistasis where the sign of the effect of each mutation depends on the allele state at the other locus. -
positive: The combined fitness effect of mutations is greater than the sum of their individual effects, often referred to as synergistic epistasis. -
negative: The combined fitness effect of mutations is less than the sum of their individual effects, often referred to as antagonistic epistasis.
Values are zero if relevant counts/instances are zero or cannot be processed.
-
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import classify_epistasis
>>> landscape = BooleanLandscape().build_from_data(
... ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> classify_epistasis(landscape, sample_cut_prob=0)["magnitude"]
1.0
Determines magnitude, sign, and reciprocal sign epistasis based on counts/estimates of motifs 19, 52, 66. Determines positive and negative epistasis by analyzing the fitness relationships within instances of these motifs.
Notes
Fractions are normalized over the directed motifs found in the constructed
graph. Missing or neutral edges can remove a square from this population.
The relation gamma_star = 1 - sign - 2*reciprocal_sign requires exact
counts of the same variable squares with no neutral effects; it is not an
identity for arbitrarily filtered or sampled graphs. See gamma_star.
Raises
-
AttributeError If landscape.graph is not an igraph.Graph object or does not exist.
-
ValueError If sample_cut_prob is invalid, or if the fitness attribute is missing.
Higher-order Epistasis¶
In addition to pairwise epistasis, higher-order epistasis that involves multiple mutations are also common in empirical fitness landscapes (e.g., Weinreich et al. 2013, Domingo et al. 2018). GraphFLA uses walsh_hadamard to estimate coefficients and summarize contributions by interaction order.
Walsh-Hadamard coefficients¶
To identify which variables interact, fitness can be represented as a sum of individual effects and interactions. For binary variables, let \(z_i=x_i-\tfrac12\), where \(x_i\in\{0,1\}\). The fitted landscape takes the form
Here, \(\varepsilon_0\) is the model's mean over uniformly weighted configurations. First-order coefficients describe background-averaged effects of individual changes, second-order coefficients describe pairwise interactions, and higher orders capture effects involving additional variables. The centered encoding determines the scaling of these coefficients.
GraphFLA implements the multistate extension of Faure et al. (2024). For a variable with \(s_i\) states, \(z_i\) is replaced by centered indicators \(\phi_{i,a}(x_i)=\mathbf{1}[x_i=a]-1/s_i\), one for each nonreference state \(a\). Interaction terms multiply indicators from distinct variables.
One call returns coefficients and an order summary. Cumulative R² and its increments describe refitted models through each order; the model variance fractions describe the highest-order fit over uniformly weighted backgrounds. Lasso provides regularized estimates. Its model spectrum and observed-data fit gains answer different questions and are reported separately.
graphfla.analysis.walsh_hadamard(
landscape,
max_order=2,
max_cells=10000000.0,
chunk_size=1000,
*,
method="ols",
alpha="cv",
cv=5,
random_state=0,
max_iter=10000,
tol=0.0001,
n_jobs=1,
) -> Dict[str, Union[pd.DataFrame, dict]]Return fitted Walsh-Hadamard coefficients and contributions by order.
Parameters
-
landscape: Landscape Built landscape with finite fitness and unique, nonmissing configurations. Use the nodes retained in
get_data(); construction filters are not reversed. All variable types, including ordinal variables, are treated as discrete states without numeric spacing. States absent from the retained data are not inferred. Fitness is used in its supplied units, without changing sign for minimization.-
max_order: int, default=2 Largest number of distinct interacting variables. Must be nonnegative. Zero fits only a constant; values above the number of variable sites include all orders. Truncation specifies a model, not evidence that omitted interactions are absent.
-
max_cells: int or float, default=1e7 Maximum number of entries in the dense design, including its constant column. Checked before enumerating terms or allocating the design. Each entry uses eight bytes. This is not a total memory limit: solver copies and workspace also require memory.
-
chunk_size: int, default=1000 Maximum number of rows in temporary feature-construction blocks. The final design is dense regardless of this value.
-
method: (ols, lasso), default="ols" Coefficient estimator. OLS raises on a rank-deficient design. Lasso minimizes squared prediction error plus an L1 coefficient penalty; it does not establish identifiability of unregularized coefficients. Neither method automatically switches to the other.
-
alpha: float or {cv}, default="cv" Positive Lasso penalty in the objective
sum((y - prediction)**2) / (2 * n_samples) + alpha * sum(abs(coef)). The constant term is not penalized and features are not standardized."cv"selects among 100 logarithmically spaced penalties from the smallest all-zero-slope penalty to one thousandth of that value, using mean validation squared error. Ignored for OLS.-
cv: int, default=5 Number of shuffled folds when
method="lasso", alpha="cv". Must be between 2 and the number of retained observations. The selected model is refitted on all observations; no held-out score is returned.-
random_state: (int, numpy.random.RandomState or None), default=0 Controls shuffled cross-validation folds. Pass an integer for reproducible splits. Used only for cross-validated Lasso.
-
max_iter: int, default=10000 Maximum coordinate-descent iterations for Lasso.
-
tol: float, default=1e-4 Positive convergence tolerance for Lasso.
-
n_jobs: int or None, default=1 Number of parallel cross-validation jobs for Lasso.
-1uses all available processors. Does not parallelize the sequence of orders.
Returns
-
result: dict Dictionary with
coefficients,order_summaryandfit_info.result["coefficients"]is a DataFrame sorted by(order, term)with:-
order: number of interacting variables. -
positions: tuple of one-based original feature positions. -
term: labelsource_position_target, joined by-for interactions. Allele-label characters%,_and-are percent-escaped; ordinary sequence labels are unchanged. -
coefficient: fitted effect in fitness units.
The order-zero row is labeled
intercept. It is the model's uniform full-space mean, not the reference configuration's fitness.result["order_summary"]has one row per order from zero through the effective maximum, with columns:-
order: maximum order in the corresponding nested fit. -
r2: cumulative training R-squared of that fit. -
delta_r2: increase over the preceding order (zero at order zero). -
rmse: training root mean squared error in fitness units. -
n_termsandrank: design width and OLS rank (NA for Lasso). -
alpha: penalty selected for that nested fit (NaN for OLS or the constant baseline; zero for a constant CV solution). -
model_variance_fraction: that order's share of the highest-order fitted model's variance over uniform product-space backgrounds. This is a model spectrum, not an observed-data R-squared increment. -
n_nonzero: number of nonzero coefficients of exactly that order in the highest-order Lasso fit (NA for OLS).
R-squared and its increments are NaN for constant fitness. Model variance fractions are NaN for a constant fitted model. Negative regularized R-squared increments are retained, not clipped.
result["fit_info"]records the estimator, dimensions, final rank and alpha, reference alleles, original column names and scoring conventions. The same metadata is attached to both tables; save it separately when exporting to formats such as CSV.-
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import walsh_hadamard
>>> landscape = BooleanLandscape().build_from_data(
... ["00", "01", "10", "11"], [10., 13., 12., 20.], verbose=False
... )
>>> result = walsh_hadamard(landscape)
>>> coefficients = result["coefficients"]
>>> round(float(coefficients.set_index("term").loc["0_1_1-0_2_1", "coefficient"]), 6)
5.0
>>> regularized = walsh_hadamard(landscape, method="lasso", alpha=0.1)
>>> regularized["fit_info"]["alpha"]
0.1
>>> result["order_summary"].order.tolist()
[0, 1, 2]
Use the multistate Walsh-Hadamard representation of Faure et al. 1.
Fit terms through max_order by ordinary least squares or explicit
Lasso regularization. Incomplete landscapes are allowed; OLS requires
a design matrix with full column rank. Build the design once and reuse it
to fit the nested models in the order summary.
Notes
Use observed allele 0 as reference for Boolean variables and the first retained observation for other variables. Construct columns of \(H^{-1}V^{-1}\) as products of centered state indicators, then fit equally weighted observations 1. Backgrounds are uniform over the Cartesian product of observed states. On a complete landscape with all orders, OLS equals the direct transform; incomplete or truncated fits estimate model coefficients.
Refit each nested model; CV selects alpha within each order using the same folds. Scores describe training fit, not held-out prediction or causal attribution. The spectrum uses exact product-space covariance, accounting for correlated multistate columns. Lasso leaves the constant unpenalized; this and the adaptive CV grid differ from the authors' analysis scripts.
References
[1] Faure, A. J., Lehner, B., Miro Pina, V., Serrano Colome, C., and Weghorn, D. (2024). An extension of the Walsh-Hadamard transform to calculate and model epistasis in genetic landscapes of arbitrary shape and complexity. PLOS Computational Biology 20(5): e1012132. https://doi.org/10.1371/journal.pcbi.1012132. Eqs. (9), (12), (23)-(25).
Raises
-
graphfla.exceptions.NotBuiltError If the landscape has not been built.
-
ValueError If inputs or parameters are invalid, the design exceeds
max_cells, OLS coefficients are not identifiable, or coefficients exceed float64.-
numpy.linalg.LinAlgError If the least-squares decomposition fails to converge.
Warns
-
UserWarning If fitness is constant and R-squared is undefined.
-
sklearn.exceptions.ConvergenceWarning If Lasso does not converge within
max_iteriterations.
Global Epistasis¶
Beyond interactions between specific pairs or sets of mutations, global epistasis describes a broader pattern where the fitness effect of a mutation systematically depends on the overall fitness of the genetic background in which it arises. This often approximately one-dimensional relationship can emerge from the combined effects of many specific, idiosyncratic interactions, which may also explain variation around the overall trend.
Common forms of these fitness-correlated trends are diminishing returns for beneficial mutations and increasing costs for deleterious mutations, with the same mutation having a less positive or more negative effect on fitter backgrounds. The returns and costs indices pool different mutations, so their signs alone do not establish mutation-specific epistasis or statistical significance.
Diminishing Returns Epistasis¶
Diminishing returns epistasis is a form of global epistasis that applies to beneficial mutations. It describes the pattern where the positive fitness effect of a mutation is smaller when it occurs in an already fit genetic background than in a less fit background. As beneficial mutations accumulate, shrinking gains can contribute to decelerating adaptation (Chou et al., 2011).
GraphFLA calculates the correlation or regression slope between background fitness and the positive fitness increase of beneficial transitions, with each retained transition contributing equally. A negative value describes diminishing gains across this pooled set. The result depends on the fitness scale and the mutations available in each background.
graphfla.analysis.diminishing_returns_index(
landscape,
method: Literal[
"pearson", "spearman", "regression"
] = "pearson",
) -> floatReturn the pooled trend of beneficial effects with background fitness.
Parameters
-
landscape: Landscape Built landscape with a directed graph of improving transitions and finite node fitness. Use a one-step neighborhood for mutation-level interpretation. The graph's retained edges determine the population; construction filters and missing configurations are not undone.
-
method: (pearson, spearman, regression), default="pearson" Pearson correlation, Spearman correlation with average ranks for ties, or the ordinary least-squares slope with an intercept. Each edge has equal weight. Effects are differences on the supplied fitness scale.
Returns
-
index: float Trend between starting fitness and the positive improvement for each retained edge. Negative values describe smaller gains on better backgrounds. Fitness is negated for minimization, so the interpretation is unchanged. Returns NaN for fewer than two eligible edges or constant starting fitness. A constant effect gives NaN for correlation and zero for regression. No significance test or p-value is returned.
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import diminishing_returns_index
>>> landscape = BooleanLandscape().build_from_data(
... ["00", "01", "10", "11"], [0., 2., 3., 4.],
... epsilon=0, verbose=False
... )
>>> round(diminishing_returns_index(landscape, method="regression"), 3)
-0.444
Notes
For each stored improving edge u -> v, correlate q(u) with q(v) - q(u),
where q = fitness for maximization and q = -fitness for minimization.
Zero-effect edges are excluded. This pools individual transitions 1,
rather than averaging effects per node or tracking a fixed mutation across
backgrounds. Edge attributes such as delta_fit are not used.
The result is descriptive: effect-sign selection, differing mutation composition (even in an additive landscape) and shared measurement error can produce a trend without demonstrating mutation-specific global epistasis. Fitness transformations can change it. Pearson and regression take O(V + E) time and O(V + B) auxiliary memory with fixed block size B; exact Spearman requires O(E) additional memory and O(E log E) time.
References
[1] Huang, M., Zhou, S. and Li, K. (2025). Augmenting Biological Fitness Prediction Benchmarks with Landscapes Features from GraphFLA. NeurIPS 38, Appendix C.3.2. https://doi.org/10.52202/085713-1180.
[2] Papkou, A. et al. (2023). A rugged yet easily navigable fitness landscape. Science 382, eadh3860, Fig. S22. https://doi.org/10.1126/science.adh3860.
Raises
-
RuntimeError If the landscape has not been built.
-
ValueError If method is invalid, the graph or fitness is missing, fitness is nonfinite, or the graph is undirected or contains worsening edges.
Warns
-
UserWarning If the requested statistic is undefined.
Increasing Cost Epistasis¶
Increasing cost epistasis is the analog of diminishing returns for deleterious mutations. It describes a pattern where a mutation's fitness cost becomes more severe in fitter backgrounds. The same mutation may have a weaker detrimental effect in less fit backgrounds (Johnson et al., 2019).
GraphFLA correlates background fitness with the positive magnitude of fitness loss, or fits a regression slope, weighting each retained deleterious transition equally. A positive value describes increasing costs. Each transition reverses a stored improving edge and starts at its better endpoint.
graphfla.analysis.increasing_costs_index(
landscape,
method: Literal[
"pearson", "spearman", "regression"
] = "pearson",
) -> floatReturn the pooled trend of deleterious cost with background fitness.
Parameters
-
landscape: Landscape Built landscape with a directed graph of improving transitions and finite node fitness. Reverse each retained edge to represent a worsening move. A mutation-level interpretation requires a reversible, one-step neighborhood. Construction filters and missing configurations remain part of the input population.
-
method: (pearson, spearman, regression), default="pearson" Pearson correlation, Spearman correlation with average ranks for ties, or the ordinary least-squares slope with an intercept. Each reverse transition has equal weight; its cost is a positive fitness difference.
Returns
-
index: float Trend between the better endpoint's fitness and the cost of moving to its worse neighbor. Positive values describe larger costs on better backgrounds. Fitness is negated for minimization. Returns NaN for fewer than two eligible edges or constant background fitness. A constant cost gives NaN for correlation and zero for regression. No significance test or p-value is returned.
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import increasing_costs_index
>>> landscape = BooleanLandscape().build_from_data(
... ["00", "01", "10", "11"], [0., 2., 3., 4.],
... epsilon=0, verbose=False
... )
>>> round(increasing_costs_index(landscape, method="regression"), 3)
-0.364
Notes
For each stored improving edge u -> v, correlate q(v) with q(v) - q(u),
using the same oriented fitness and edge population as
diminishing_returns_index. Zero-effect edges are excluded. This is the
pooled cost convention of 1; it is not the distribution of
mutation-specific regressions studied by Johnson et al. 2. Its sign
alone does not establish global epistasis or statistical significance.
References
[1] Huang, M., Zhou, S. and Li, K. (2025). Augmenting Biological Fitness Prediction Benchmarks with Landscapes Features from GraphFLA. NeurIPS 38, Appendix C.3.2. https://doi.org/10.52202/085713-1180.
[2] Johnson, M. S. et al. (2019). Higher-fitness yeast genotypes are less robust to deleterious mutations. Science 366, 490-493, Figs. 3-4. https://doi.org/10.1126/science.aay4199.
Raises
-
RuntimeError If the landscape has not been built.
-
ValueError If method is invalid, the graph or fitness is missing, fitness is nonfinite, or the graph is undirected or contains worsening edges.
Warns
-
UserWarning If the requested statistic is undefined.
Gamma Statistic¶
Gamma¶
The gamma statistic (\(\gamma\)) measures the amount of epistasis in a fitness landscape through the correlation between fitness effects of mutations across different genetic backgrounds (Ferretti et al., 2016). It compares the effects of the same mutation in backgrounds differing by a single mutation at another locus. In simpler terms, it quantifies how the effect of a particular mutation is altered by another mutation in the genetic background, averaged across the landscape. For individual two-locus, two-allele squares, the following relationships hold; the ranges overlap and do not define classification thresholds for an entire landscape:
- \(\gamma = 1\): no epistasis; mutation effects are unchanged across backgrounds.
- Magnitude epistasis \(\rightarrow 0 \leq \gamma < 1\): mutation effects retain their signs but differ in magnitude, giving a nonnegative correlation.
- Sign epistasis \(\rightarrow -\frac{1}{3} \leq \gamma < 1\): the effect of one mutation changes sign, and the correlation can be positive, zero, or negative.
- Reciprocal sign epistasis \(\rightarrow -1 \leq \gamma < 0\): the effects of both mutations change sign, giving a negative correlation.
graphfla.analysis.gamma(landscape, n_jobs=-1) -> floatReturn the correlation of mutation effects across neighboring backgrounds.
Parameters
-
landscape: Landscape Built landscape. Fitness is used on its supplied scale; apply any scientifically appropriate log transformation before construction. All observed allele pairs at two distinct variables are considered. Retained configurations define the population, independently of graph edges, ordinal step restrictions and construction epsilon.
-
n_jobs: int or None, default=-1 Number of joblib workers.
-1uses all available CPUs and1runs serially.Nonefollows the active joblib configuration, or uses one worker without a configuration. Zero is invalid.
Returns
-
gamma_value: float Non-centered correlation in [-1, 1]. One means equal mutation effects across every observed square; negative values indicate opposing effects. Returns NaN if there are no complete squares or every effect on those squares is zero.
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import gamma
>>> landscape = BooleanLandscape().build_from_data(
... ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> round(gamma(landscape, n_jobs=1), 6)
0.888889
Notes
For parallel effects a and b, pool their products and squared effects:
sum(a*b) / sum((a*a + b*b)/2) over both directions of every complete
two-variable square 1. Each allele-pair combination has equal weight;
averaging square ratios gives a different statistic. Fitness offsets and
nonzero linear rescaling leave gamma unchanged.
Missing corners exclude a square from both sums. This describes the observed squares, not the paper's separate distance-correlation estimator for missing data (Eq. (2)). A value of one on incomplete data does not establish global additivity.
References
[1] Ferretti, L. et al. (2016). Measuring epistasis in fitness landscapes: The correlation of fitness effects of mutations. Journal of Theoretical Biology, 396, 132-143. Eqs. (1), (3), Appendix C.1. https://doi.org/10.1016/j.jtbi.2016.01.037
Raises
-
graphfla.exceptions.NotBuiltError If the landscape has not been built.
-
ValueError If the graph lacks fitness values, or the worker count is invalid.
Warns
-
UserWarning If fewer than two variables remain in the built landscape.
Gamma Star¶
The gamma-star statistic (\(\gamma^*\)) is a variant of \(\gamma\) that focuses only on sign consistency, ignoring the magnitude of fitness effects. It indicates whether mutations tend to have consistent directional effects across different genetic backgrounds.
graphfla.analysis.gamma_star(landscape, n_jobs=-1) -> floatReturn the correlation of mutation-effect signs across backgrounds.
Parameters
-
landscape: Landscape Built landscape. Positive, zero and negative fitness differences are assigned +1, 0 and -1, respectively. Only exact ties are neutral; construction epsilon does not set a sign tolerance for this metric.
-
n_jobs: int or None, default=-1 Number of joblib workers.
-1uses all available CPUs and1runs serially.Nonefollows the active joblib configuration, or uses one worker without a configuration. Zero is invalid.
Returns
-
gamma_star_value: float Sign correlation in [-1, 1]. One indicates consistent nonzero signs, minus one indicates reversed signs, and zero indicates cancellation or absence of nonzero parallel products. Returns NaN if there are no complete squares or all effects on those squares are neutral.
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import gamma_star
>>> landscape = BooleanLandscape().build_from_data(
... ["00", "01", "10", "11"], [0, 0, 1, 2], verbose=False)
>>> round(gamma_star(landscape, n_jobs=1), 6)
0.666667
Notes
Apply the pooling formula of gamma to effect signs 1. A zero
effect contributes zero to the numerator and its squared-effect term;
it does not cause the entire square to be discarded. The function uses
Eq. (11) with zero tolerance, not the optional positive-tolerance variant.
The identity gamma_star = 1 - phi_sign - 2*phi_reciprocal (Eq. (12))
requires no neutral effects and the same square population. It need not
hold for classify_epistasis when graph filtering removes edges,
graph motifs differ from variable squares, or motif counts are sampled.
References
[1] Ferretti, L. et al. (2016). Measuring epistasis in fitness landscapes: The correlation of fitness effects of mutations. Journal of Theoretical Biology, 396, 132-143. Eqs. (10)-(12), Appendix C.3. https://doi.org/10.1016/j.jtbi.2016.01.037
Raises
-
graphfla.exceptions.NotBuiltError If the landscape has not been built.
-
ValueError If the graph lacks fitness values, or the worker count is invalid.
Warns
-
UserWarning If fewer than two variables remain in the built landscape.
Idiosyncratic Epistasis¶
The fitness-correlated trends described above summarize how mutation effects change with background fitness. Individual mutations can also depend on the specific alleles present in a background. For example, a substitution that improves a protein's activity in one sequence may become harmful after another residue changes. Such interactions are described as idiosyncratic epistasis and help explain why a mutation's effect in one genotype may not predict its effect in another.
Lyons et al. (2020) introduced the idiosyncratic index to quantify background dependence. For a given mutation, it compares the standard deviation of its effects across matching backgrounds with that of fitness differences in an equally sized sample of random genotype pairs. GraphFLA summarizes the landscape by averaging these ratios across eligible directed mutations, giving each mutation equal weight.
When defined, an index of zero indicates constant mutation effects across the observed backgrounds. A value near one indicates variation comparable to random genotype-pair differences; estimates can exceed one. The index measures variation, rather than the fraction of epistatic mutations. Because a nonlinear global fitness relationship can also generate background dependence, a positive index alone does not distinguish specific interactions from global epistasis.
Per-mutation Index¶
graphfla.analysis.idiosyncratic_index(
landscape, mutation, min_pairs: int = 3, *, seed=None
) -> floatReturn a mutation's idiosyncratic index from matched backgrounds.
Parameters
-
landscape: Landscape Built landscape with unique configurations and finite fitness values. Both matched backgrounds and random controls use the genotypes retained in
landscape.get_data().-
mutation: tuple of (source, position, target) Allele substitution to evaluate.
positionis a configuration-column label fromlandscape.data_types, not a positional column index.sourceandtargetmust be distinct observed alleles at that position. For example,(0, "bit_0", 1)changes the first Boolean feature from 0 to 1.-
min_pairs: int, default=3 Minimum number of observed matching backgrounds, at least 2. All matching backgrounds are used when this threshold is met. The control then contains the same number of random pairs. This threshold is a GraphFLA estimation guard, not a cutoff specified in 1.
-
seed: int or None, default=None Seed for a local NumPy RandomState. An integer reproduces the control sample for the same ordered input; None starts a fresh random stream. NumPy's global RNG state is not modified.
Returns
-
index: float Ratio of observed-effect SD to sampled-control SD. Values can exceed 1. Returns NaN for a constant-fitness landscape, fewer than
min_pairsbackgrounds, or a sampled control with zero SD. A well-defined value of zero indicates constant mutation effects across the observed backgrounds.
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import idiosyncratic_index
>>> sequences = ["000", "001", "010", "011", "100", "101", "110", "111"]
>>> fitness = [0., 1., 2., 3., 1., 2., 4., 6.]
>>> landscape = BooleanLandscape().build_from_data(
... sequences, fitness, epsilon=0, verbose=False
... )
>>> value = idiosyncratic_index(landscape, (0, "bit_0", 1), seed=0)
>>> round(value, 3)
0.308
The index compares the standard deviation of one mutation's effects with that of an equally sized sample of random genotype-pair differences 1. Effects and controls use the fitness scale supplied in the landscape.
Notes
Match genotypes at all positions except the focal one. Divide the SD of
their fitness differences by that of equally many random-pair differences,
sampling both endpoints independently with replacement. Both SDs use
ddof=0; matching does not depend on graph edges.
With seed=None, results can differ across calls. Fitness is used as
supplied. The index measures background dependence, including
variation that can arise from a nonlinear global fitness map.
References
[1] Lyons, Daniel M., Zhengting Zou, Haiqing Xu, and Jianzhi Zhang. "Idiosyncratic epistasis creates universals in mutational effects and evolutionary trajectories." Nature Ecology & Evolution 4 (2020): 1685-1693. https://doi.org/10.1038/s41559-020-01286-y.
Raises
-
graphfla.exceptions.NotBuiltError If the landscape has not been built.
-
ValueError If
min_pairsor seed is invalid, the mutation uses an unknown position or allele, or source and target are equal. Also raised for missing configuration values, duplicate configurations, or nonfinite fitness.
Warns
-
RuntimeWarning If the sampled control has zero SD. No replacement sample is drawn.
Global (Landscape-Wide) Index¶
graphfla.analysis.global_idiosyncratic_index(
landscape, n_jobs=-1, seed=None, min_pairs: int = 3
) -> floatReturn the mean idiosyncratic index across directed mutations.
Parameters
-
landscape: Landscape Built landscape with unique configurations and finite fitness values. Mutation effects and random controls both use the genotypes retained in
landscape.get_data(). Apply the intended fitness transformation and population selection before constructing the landscape.-
n_jobs: int or None, default=-1 Number of parallel jobs for background matching. -1 uses all available cores; 1 runs serially. This does not change a fixed-seed result.
-
seed: int or None, default=None Seed for a local NumPy RandomState. An integer reproduces the result for the same ordered input and parameters. None starts a fresh random stream. NumPy's global RNG state is not modified.
-
min_pairs: int, default=3 Minimum number of matching backgrounds for a mutation to contribute, at least 2. Mutations below this threshold are omitted. All matching backgrounds of eligible mutations are used, with equally sized random controls. The paper specifies no minimum for this index.
Returns
-
mean_index: float Arithmetic mean of the eligible mutation indices. Values can exceed 1. Returns NaN for constant fitness, no eligible mutations, or a zero-SD control for any eligible mutation. A failed control is not omitted from the mean or resampled.
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import global_idiosyncratic_index
>>> sequences = ["000", "001", "010", "011", "100", "101", "110", "111"]
>>> fitness = [0., 1., 2., 3., 1., 2., 4., 6.]
>>> landscape = BooleanLandscape().build_from_data(
... sequences, fitness, epsilon=0, verbose=False
... )
>>> round(global_idiosyncratic_index(landscape, n_jobs=1, seed=0), 3)
0.416
Apply the SD ratio defined by Lyons et al. 1 to each eligible directed mutation and return their arithmetic mean. Every mutation receives equal weight, irrespective of its number of observed backgrounds.
Notes
Apply idiosyncratic_index's calculation to each ordered allele pair
at each position. Forward and reverse mutations receive independent
controls. Average over eligible mutations, not positions or backgrounds.
Matching uses configuration values rather than graph edges. Controls use only retained genotypes, so graph pruning changes the reference population.
References
[1] Lyons, Daniel M., Zhengting Zou, Haiqing Xu, and Jianzhi Zhang. "Idiosyncratic epistasis creates universals in mutational effects and evolutionary trajectories." Nature Ecology & Evolution 4 (2020): 1685-1693. https://doi.org/10.1038/s41559-020-01286-y.
Raises
-
graphfla.exceptions.NotBuiltError If the landscape has not been built.
-
ValueError If
min_pairsor the random seed is invalid, or configurations are missing or duplicated, or fitness contains nonfinite values.
Warns
-
RuntimeWarning If any eligible mutation's sampled control has zero SD.
Extradimensional Bypass Analysis¶
Detects extradimensional bypasses around reciprocal-sign-epistasis (RSE) motifs.
Description
Reciprocal sign epistasis occurs when the wildtype (ab) and the double mutant (AB) are both fitter than the two intermediate single mutants (aB, Ab). This creates a fitness valley along the direct path from ab to AB that natural selection cannot cross. However, when the landscape has more than two relevant dimensions, the population can sometimes traverse a third-position intermediate (i.e., gain and later lose an extra mutation) to bypass the valley. Such indirect uphill paths are called extradimensional bypasses and they greatly affect a landscape's effective navigability.
For each RSE motif in the landscape (igraph isomorphism class 19), this function checks whether the directed improving graph admits any path from ab to AB. The fraction of RSE motifs that do admit such a bypass is the bypass proportion.
graphfla.analysis.extradimensional_bypass(
landscape,
sample_cut_prob="auto",
seed=None,
time_budget=15.0,
) -> Dict[str, Union[float, int]]Return a summary of bypasses around reciprocal-sign epistasis motifs.
Parameters
-
landscape: Landscape The fitness landscape object, containing landscape.graph as an igraph.Graph with a "fitness" vertex attribute.
-
sample_cut_prob: auto, default="auto" Controls the 4-motif search.
"auto"(default) picks a pruning probability from a runtime ladder targetingtime_budget(small landscapes resolve to exact). This is not a hard time limit.0orNoneforces exact enumeration; a float in(0, 1]sets the igraph pruning probability directly (higher -> faster, less accurate).-
seed: int or None, default=None Seed for the sampling RNG, making approximate results reproducible.
-
time_budget: float, default=15.0 Target wall-clock seconds for
sample_cut_prob="auto".
Returns
-
result: dict Summary with the following keys:
-
bypass_proportion: proportion of reciprocal-sign-epistasis motifs for which an extradimensional bypass exists (float between 0 and 1). -
average_bypass_length: average length of extradimensional bypasses for motifs where such bypasses exist; NaN if none exist. -
total_motifs: integer total number of type-19 motifs analyzed. -
motifs_with_bypass: integer number of motifs that have a bypass.
-
Examples
>>> from graphfla.landscape import BooleanLandscape
>>> from graphfla.analysis import extradimensional_bypass
>>> landscape = BooleanLandscape().build_from_data(
... ["00", "01", "10", "11"], [0, 1, 2, 4], verbose=False)
>>> extradimensional_bypass(landscape, sample_cut_prob=0)["bypass_proportion"]
0.0
For each motif representing reciprocal sign epistasis (type 19), this function identifies whether accessible evolutionary paths exist that bypass the direct path between the double mutant nodes. Such indirect paths are called extradimensional bypasses and allow evolution to traverse fitness valleys that would otherwise be inaccessible under strong selection.
Notes
Reciprocal sign epistasis occurs when both the wildtype (ab) and double mutant (AB) have higher fitness than both single mutants (aB, Ab). This creates a fitness valley that prevents direct evolutionary access between ab and AB. Extradimensional bypasses are indirect paths through the broader fitness landscape that circumvent this valley.
Raises
-
AttributeError If landscape.graph is not an igraph.Graph object or does not exist.
-
ValueError If sample_cut_prob is not between 0 and 1, or if fitness attribute missing.