Landscape ¶
Landscape(
kind: str = "default",
maximize: bool = True,
input_handler: Optional[InputHandler] = None,
neighbor_generator: Optional[NeighborGenerator] = None,
strategy_key: Optional[str] = None,
)Class implementing the fitness landscape object.
This class provides a foundational structure for fitness landscapes, conceptualized as a mapping from a genotype (configuration) space to a fitness value. It typically represents the landscape as a directed graph where nodes are genotypes and edges connect mutational neighbors, pointing towards fitter variants.
The landscape can be either constructed from raw data (using build_from_data()) or
leveraging existing graph (using build_from_graph()).
After construction:
-
The graph object representing the landscape can be accessed via
Landscape.graph. -
Tabular information about the configurations can be obtained with
get_data(). -
Basic landscape properties are available via
describe(). -
Other methods in the
graphfla.analysisandgraphfla.plottingmodules can be used for advanced analysis and visualization.
Methods at a glance
get_params()- Return the constructor parameters as a dict.
register_input_handler()- Register a custom input handler for a data type on this instance.
register_neighbor_generator()- Register a custom neighbor generator for a data type on this instance.
build_from_data()- Construct the landscape graph and properties from configuration data.
get_data()- Extracts landscape data as a pandas DataFrame.
get_lon()- Constructs and returns the Local Optima Network (LON).
describe()- Return a structured summary of the landscape as a dict.
build_from_graph()- Construct a landscape from a saved graph file.
to_graph()- Save the landscape graph and essential attributes to a file.
Parameters
-
type: str, default='default' The type of landscape to create. This determines the input-preparation and neighbor generation strategies used. Options include 'boolean', 'dna', 'rna', 'protein', or 'default' for general landscapes.
-
maximize: bool, default=True Determines the optimization direction. If True, the landscape seeks higher fitness values (peaks are optima). If False, it seeks lower fitness values (valleys are optima).
Attributes
-
graph: ig.Graph or None The directed ig.Graph representing the fitness landscape. Nodes represent configurations (genotypes) and edges connect neighboring configurations, typically pointing from lower to higher fitness if
maximizeis True. Each node usually has a 'fitness' attribute. Populated after callingbuild_from_dataorbuild_from_graph. Fitness difference between neighboring nodes is stored in the edge attribute 'delta_fit'.-
configs: pandas.Series or None A pandas Series mapping node indices (int) to their corresponding configuration representation (often a tuple). This represents the genotypes in the landscape. Populated after calling
build_from_dataor when loading a graph viabuild_from_graphif the file contains configs data.-
config_dict: dict or None A dictionary describing the encoding scheme for configuration variables. Keys are typically integer indices of variables, and values are dictionaries specifying properties like 'type' (e.g., 'boolean', 'categorical') and 'max' (maximum encoded value). Populated after calling
build_from_data.-
data_types: dict or None A dictionary specifying the data type for each variable in the configuration space (e.g., {'var_0': 'boolean', 'var_1': 'categorical'}). Validated and stored during
build_from_data. Required for certain distance calculations.-
n_configs: int or None The total number of configurations (nodes) in the landscape graph. Populated after calling
build_from_dataorbuild_from_graph.-
n_vars: int or None The number of variables (dimensions) defining a configuration in the genotype space. Populated after calling
build_from_dataor inferred bybuild_from_graph.-
n_edges: int or None The total number of directed edges (connections) in the landscape graph. Populated after calling
build_from_dataorbuild_from_graph.-
n_lo: int or None The number of distinct local optima (plateau-aware): each neutral plateau-LO counts once, plus every single-point LO. This is the "number of local optima". Populated after graph analysis.
-
n_lo_members: int or None The total number of local-optimum member nodes (every member of every plateau-LO plus single-point LOs); equals
len(lo_index). For a landscape with no neutral plateaus this equalsn_lo. Populated after graph analysis.-
lo_index: list[int] or None A sorted list of all node indices that are local optima (plateau-aware); its length is
n_lo_members. Populated after graph analysis.-
go_index: int or None The node index of the global optimum (the configuration with the highest or lowest fitness). Populated after graph analysis.
-
go: dict or None A dictionary containing the attributes (including 'fitness') of the global optimum node. Populated after graph analysis.
-
lon: ig.graph or None The Local Optima Network (LON) graph, if calculated via
get_lon. Nodes in the LON are local optima from the main landscape graph, and edges represent accessibility between their basins (Ochoa 2021).-
has_lon: bool Flag indicating whether the LON has been calculated and stored in the
lonattribute.-
maximize: bool Indicates whether the objective is to maximize (True) or minimize (False) the fitness values. Set during
build_from_dataorbuild_from_graph.-
epsilon: float Neutrality threshold. Neighbors with
|fitness_a - fitness_b| <= epsilonare classified as neutral rather than improving/worsening. Whenepsilon > 0, a plateau layer is constructed and downstream analyses become plateau-aware. Defaults to 0.-
plateaus: dict or None Mapping from 0-based plateau ID to the list of member node indices. Only multi-member plateaus are stored; singleton nodes have
plateau_id = -1. Populated whenepsilon > 0and neutral pairs exist.-
n_plateau: int or None The number of neutral plateaus (multi-member connected components of neutral neighbors). Populated when
epsilon > 0and neutral pairs exist; 0 when no plateaus; None before construction.-
plateau_lo_index: list[int] or None Plateau IDs (0-based) whose plateaus are local optima. Does not include single-point LOs.
-
n_plateau_lo: int or None len(plateau_lo_index).-
verbose: bool The verbosity level set during initialization or construction.
-
_is_built: bool Internal flag indicating if the landscape has been populated via
build_from_dataorbuild_from_graph.
n_configs
¶
Number of configurations (nodes). Derived from the graph once built
(single source of truth); falls back to the working count during
construction, before self.graph exists.
n_edges
¶
Number of directed edges. Derived from the graph once built (single source of truth); falls back to the working count during construction.
shape
¶
Return the shape (n_configs, n_edges) of the landscape graph.
configs
: Optional[pd.Series]
¶
Per-node configuration tuple Series (built lazily, then cached).
The numeric _configs_array is the source of truth and is what graph
construction consumes; this tuple Series is a downstream artifact
used only by analyses that need per-node configuration tuples. It is
materialised from _configs_array on first access -- builds that never
read configs (the common construction-only path) skip the cost
entirely. Returns None only when neither the cached Series nor a
numeric array is available (e.g. an unbuilt landscape, or a graph load
from which configurations could not be reconstructed).
Note: a landscape.configs is None check materialises the Series
if _configs_array is present.
basins
: pd.Series
¶
Per-node greedy basin size (size_basin_greedy), computed lazily.
accessible_paths
: pd.Series
¶
Per-node accessible-basin size (size_basin_accessible), computed lazily.
dist_to_go
: pd.Series
¶
Per-node configuration distance to the nearest global optimum
(dist_go), computed lazily.
neighbor_fitness
: pd.Series
¶
Per-node mean neighbour fitness (mean_neighbor_fit), computed lazily.
pagerank
: pd.Series
¶
Per-node PageRank centrality, computed lazily.
PageRank is not used by any landscape metric; it is only an optional,
descriptive node attribute. To keep it off the construction critical
path (where it was the single dominant cost) it is computed on first
access here -- producing values identical to the eager computation
(weighted by delta_fit when present, directed=True).
build_from_data()
build_from_data(
X: Any,
f: Union[Series, list, ndarray],
data_types: Optional[Dict[str, str]] = None,
epsilon: float = 0,
tau: Optional[float] = None,
filter_mode: str = "any",
n_edit: int = 1,
neighborhood_strategy: Optional[str] = None,
verbose: Optional[bool] = True,
) -> "Landscape"Construct the landscape graph and properties from configuration data.
Parameters
-
X: pandas.DataFrame or numpy.ndarray or a list of strings The configuration data, where each row represents a genotype or configuration, and columns represent variables or sites. Column names are converted to strings and must remain unique. Names used for landscape attributes, such as 'fitness', 'out_degree' and 'is_lo', are reserved.
-
f: pandas.Series, list, or numpy.ndarray The fitness values corresponding to each configuration in
X. Must have the same length asX. Values are normalized to float64 before fitness arithmetic.-
data_types: dict[str, str], default=None A dictionary specifying the type of each variable (column in
Xif DataFrame, or inferred column index if ndarray). Keys must match column names/indices, and values must be one of 'boolean', 'categorical', or 'ordinal'. This information is crucial for determining neighborhood relationships and calculating distances. Optional: the typed landscape subclasses (BooleanLandscape,OrdinalLandscape,DNALandscape/RNALandscape/ProteinLandscape) auto-detect it. Supply it explicitly for a genericLandscape(kind="default")with heterogeneous (mixed) columns.-
epsilon: float, default=0 Neutrality threshold. Neighboring configurations whose absolute fitness difference is
<= epsilonare treated as neutral (equal-fitness) rather than strictly improving/worsening. Whenepsilon > 0, a plateau layer is constructed: neutral neighbors are grouped into plateaus via connected components, and downstream analyses (local optima, basins, accessible paths) become plateau-aware. A higher epsilon produces a smoother landscape with fewer, larger local optima;epsilon=0preserves strict-inequality behavior (the default).-
tau: float, default=None Functional threshold. Configurations whose fitness is above (when
maximize=True) or below (whenmaximize=False) this value are considered "functional". Used together withfilter_modeto focus the landscape on biologically or practically relevant regions.-
filter_mode: str, default='any' How to apply the functional threshold
tau. Options:-
'any': Pre-construction filter. Remove non-functional configurations before building the graph. If maximize: keep fitness >= tau; if minimize: keep fitness <= tau. -
'both': Post-construction filter. Keep all configurations, but remove edges where both endpoints are non-functional (both < tau when maximize, both > tau when minimize). Preserves transitions between functional and non-functional regions.
-
-
n_edit: int, default=1 The edit distance defining the neighborhood. For
pairwiseandbroadcaststrategies, an undirected neighbor pair is kept when Hamming distance is positive and<= n_edit. Theactivestrategy uses type-specific generators, which only supportn_edit=1for boolean, sequence, and default landscapes (usepairwiseorbroadcastfor multi-edit Hamming graphs).-
neighborhood_strategy: str, default=None Strategy for identifying neighboring configurations. When
None(the default), the class default is used (_default_neighborhood_strategy—'auto'for most landscapes,'active'for ordinal). On ordinal (or mixed landscapes containing an ordinal variable),'auto'resolves to'active'and the Hamming-based'pairwise'/'broadcast'strategies emit a warning, because they ignore the ±1-step ordinal adjacency. Options:-
'auto': Automatically selects the fastest strategy based on dataset size, sequence length, and alphabet size. Uses'pairwise'when the full distance matrix fits in memory (~4 GiB), falls back to'broadcast'when vectorized per-config distances are cheaper than candidate generation, and defaults to'active'otherwise. -
'active': For each configuration, enumerates all possible single-edit (orn_edit-edit) mutant neighbors and checks whether they exist in the dataset via hash lookup. Efficient for dense datasets where most proposed neighbors are present. -
'pairwise': Computes the full pairwise Hamming distance matrix usingscipy.spatial.distance.pdist. Very fast for small-to-moderate datasets (up to ~25 000 configurations) thanks to highly optimized C code. Memory usage scales asO(n_configs^2). -
'broadcast': For each configuration, computes its Hamming distance to all other configurations using vectorized NumPy operations. Suitable for large datasets with long sequences and sparse sampling, where the'active'strategy would waste time generating candidates that don't exist and'pairwise'would exceed memory limits.
-
-
verbose: bool, default=True If provided, overrides the instance's verbosity setting for this method call. Controls printed output during the build process.
Returns
-
Landscape The populated instance itself (
self), to support method chaining (e.g.BooleanLandscape().build_from_data(X, f)).
This method takes genotype-phenotype data (configurations X and their
corresponding fitness values f) and builds the underlying graph
structure of the fitness landscape. Nodes represent configurations, and
edges connect neighbors based on the specified edit distance (n_edit).
It then determines the basic landscape properties (local optima and the
global optimum). More expensive analyses — basins of attraction,
accessible paths, distance-to-optimum and neighbour fitness — are computed
lazily on first access via the .basins / .accessible_paths /
.dist_to_go / .neighbor_fitness properties.
This method populates the core attributes of the Landscape instance.
Notes
The construction process involves several steps:
-
Resolve the type-specific preparation and neighbor-generation strategies.
-
Apply the optional pre-construction fitness filter.
-
Prepare the raw input based on the landscape type.
-
Warn on missing values (no automatic row drops); drop duplicate configurations; encode step errors if NaNs remain.
-
Encode the data and cache configuration metadata.
-
Construct the core directed landscape graph.
-
Apply post-construction pruning and remap cached metadata.
-
Build neutral plateaus and analyze derived landscape properties.
Raises
-
RuntimeError If the
build_from_dataorbuild_from_graphmethod has already been called on this instance. Create a new instance to rebuild.-
ValueError If input data
X,f, ordata_typesare invalid (e.g., mismatched lengths, empty data, invalid types indata_types), or if the construction process fails (e.g., resulting graph is empty).-
TypeError If input types are incorrect (e.g.,
data_typesis not a dict).
get_data()
get_data(
lo_only: bool = False, include_pagerank: bool = False
) -> pd.DataFrameExtracts landscape data as a pandas DataFrame.
Parameters
-
lo_only: bool, default=False If True, returns data only for the configurations identified as local optima. If the Local Optima Network (LON) has been computed (see
get_lon), data from the LON graph is returned. Otherwise, it returns data from the main graph filtered for local optima nodes. If False, returns data for all configurations in the main graph.-
include_pagerank: bool, default=False If True, compute (if needed) and include a
pagerankcolumn. PageRank is otherwise omitted -- it is an optional, comparatively expensive centrality that most callers do not need, soget_datano longer triggers it as a hidden side effect.
Returns
-
pandas.DataFrame A DataFrame containing the attributes of the landscape nodes. Index corresponds to the node indices.
Returns a DataFrame where rows correspond to configurations (nodes) and columns correspond to their attributes (e.g., fitness, degree, basin information, original features). Feature columns appear in the original input order.
Raises
-
RuntimeError If the landscape has not been built (via
build_from_dataorbuild_from_graph) before calling this method. Iflo_only=Trueand the LON graph (self.lon) is unexpectedly None despiteself.has_lonbeing True.
get_lon()
get_lon(
mlon: bool = True,
min_edge_freq: int = 3,
trim: Optional[int] = None,
verbose: Optional[bool] = None,
) -> ig.GraphConstructs and returns the Local Optima Network (LON).
Parameters
-
mlon: bool, default=True If True, also build the monotonic LON (edges restricted to non-worsening transitions).
-
min_edge_freq: int, default=3 Keep a LON edge only when the number of basin transitions between two optima is strictly greater than this threshold.
-
trim: int, default=None If given, keep only the
trimstrongest outgoing edges per node.-
verbose: bool, default=None Verbosity override; defaults to the landscape's own
verbose.
Returns
-
ig.Graph The constructed Local Optima Network graph. The graph is also stored in the
self.lonattribute, andself.has_lonis set to True.
The LON is a coarse-grained representation of the fitness landscape where nodes are the local optima of the original landscape, and edges represent the possibility of transitions between their basins of attraction, typically weighted by the fitness difference or distance between the optima. This method requires the landscape graph to be built and local optima to be identified.
The landscape's own graph, configurations, optima and config_dict
are supplied automatically; the parameters below control the LON itself
and are forwarded to graphfla.lon.get_lon.
Raises
-
RuntimeError If the landscape has not been built, or if essential attributes (
graph,configs,lo_index,config_dict) required for LON construction are missing.
describe()
describe() -> Dict[str, Any]Return a structured summary of the landscape as a dict.
Returns
-
dict class,kind,built,maximizeandepsilonare always present. When built, size/optima fields (n_vars,n_configs,n_edges,n_lo,go_index), the calculation flags, and -- if a plateau layer exists -- the plateau counts are added.
Unlike a printed summary, the returned mapping is composable and
testable -- callers can log it, assert on it, or render it. For a quick
human-readable view use print(landscape) (see __str__).
build_from_graph()
build_from_graph(
filepath: str, *, verbose: bool = True
) -> "Landscape"Construct a landscape from a saved graph file.
Parameters
-
filepath: str Path to the saved graph file (.graphml).
-
verbose: bool, default=True Controls verbosity of output during loading and analysis.
Returns
-
Landscape A new instance populated with the graph and inferred properties.
This class method creates a new landscape instance by loading a previously saved graph, avoiding the need to reconstruct the landscape from original configuration data. This is significantly faster than building from scratch.
Notes
This method will:
-
Load the saved graph structure and attributes
-
Infer essential landscape properties from the graph
-
Recalculate local optima and global optimum from the graph structure
Previously computed attributes (basins, accessible paths, distances, neighbor fitness) are preserved from the saved graph if present.
Some specialized attributes from subclasses (like sequence_length in SequenceLandscape) will be inferred where possible.
Only load GraphML from trusted sources: embedded configuration metadata
is parsed with ast.literal_eval (safe against arbitrary code
execution, but not a substitute for validating untrusted files).
Raises
-
ValueError If the file cannot be read or doesn't contain valid graph data.
-
FileNotFoundError If the specified file doesn't exist.
to_graph()
to_graph(filepath: str) -> NoneSave the landscape graph and essential attributes to a file.
Parameters
-
filepath: str The path where the graph file will be saved. If the file doesn't end with '.graphml', this extension will be added automatically.
This method serializes the landscape's graph structure and relevant
attributes to a GraphML file, which can later be loaded using build_from_graph.
This allows efficient storage and sharing of landscapes without requiring
re-construction from scratch.
Notes
The GraphML format preserves the graph structure and all vertex/edge attributes.
In addition to the graph itself, essential landscape attributes like maximize
and epsilon are stored as graph attributes.
Raises
-
NotBuiltError If the landscape has not been built.
-
ValueError If the graph cannot be saved to the specified path.
References
[Wright 1932] Wright, S. The roles of mutation, inbreeding, crossbreeding and selection in evolution. Proceedings Sixth International Congress Genetics 1, 356-366 (1932).
[Papkou 2023] Papkou, A. et al. A rugged yet easily navigable fitness landscape. Science 382, eadh3860 (2023).
[Li 2016] Li, C. et al. The fitness landscape of a tRNA gene. Science 352, 837-840 (2016).
[Puchta 2016] Puchta, O. et al. Network of epistatic interactions within a yeast snoRNA. Science 352, 840-844 (2016).
[Poelwijk 2007] Poelwijk, F. J. et al. Empirical fitness landscapes reveal accessible evolutionary paths. Nature 445, 383-386 (2007).
[Carneiro 2010] Carneiro, M. & Hartl, D. L. Adaptive landscapes and protein evolution. PNAS 107, 1747-1751 (2010).
[Ochoa 2021] Ochoa, G. et al. Local optima networks: A survey. Journal of Heuristics 27, 79-134 (2021).
Examples
>>> import pandas as pd
>>> import numpy as np
>>> X_data = pd.DataFrame({'var_0': [0, 0, 1, 1], 'var_1': [0, 1, 0, 1]})
>>> f_data = pd.Series([1.0, 2.0, 3.0, 2.5])
>>> landscape = BooleanLandscape().build_from_data(X_data, f_data, verbose=False)
>>> landscape.describe()["n_lo"]
1
>>> repr(landscape)
'BooleanLandscape(maximize=True)'
>>> print(f"Number of configurations: {landscape.n_configs}")
Number of configurations: 4
>>> print(f"Global optimum fitness: {landscape.go['fitness']}")
Global optimum fitness: 3.0