Landscape

Landscape(
    kind: str = "default",
    maximize: bool = True,
    input_handler: Optional[InputHandler] = None,
    neighbor_generator: Optional[NeighborGenerator] = None,
    strategy_key: Optional[str] = None,
)

Class implementing the fitness landscape object.

This class provides a foundational structure for fitness landscapes, conceptualized as a mapping from a genotype (configuration) space to a fitness value. It typically represents the landscape as a directed graph where nodes are genotypes and edges connect mutational neighbors, pointing towards fitter variants.

The landscape can be either constructed from raw data (using build_from_data()) or leveraging existing graph (using build_from_graph()).

After construction:

  • The graph object representing the landscape can be accessed via Landscape.graph.

  • Tabular information about the configurations can be obtained with get_data().

  • Basic landscape properties are available via describe().

  • Other methods in the graphfla.analysis and graphfla.plotting modules can be used for advanced analysis and visualization.

Methods at a glance

get_params()
Return the constructor parameters as a dict.
register_input_handler()
Register a custom input handler for a data type on this instance.
register_neighbor_generator()
Register a custom neighbor generator for a data type on this instance.
build_from_data()
Construct the landscape graph and properties from configuration data.
get_data()
Extracts landscape data as a pandas DataFrame.
get_lon()
Constructs and returns the Local Optima Network (LON).
describe()
Return a structured summary of the landscape as a dict.
build_from_graph()
Construct a landscape from a saved graph file.
to_graph()
Save the landscape graph and essential attributes to a file.

Parameters

type : str, default='default'

The type of landscape to create. This determines the input-preparation and neighbor generation strategies used. Options include 'boolean', 'dna', 'rna', 'protein', or 'default' for general landscapes.

maximize : bool, default=True

Determines the optimization direction. If True, the landscape seeks higher fitness values (peaks are optima). If False, it seeks lower fitness values (valleys are optima).

Attributes

graph : ig.Graph or None

The directed ig.Graph representing the fitness landscape. Nodes represent configurations (genotypes) and edges connect neighboring configurations, typically pointing from lower to higher fitness if maximize is True. Each node usually has a 'fitness' attribute. Populated after calling build_from_data or build_from_graph. Fitness difference between neighboring nodes is stored in the edge attribute 'delta_fit'.

configs : pandas.Series or None

A pandas Series mapping node indices (int) to their corresponding configuration representation (often a tuple). This represents the genotypes in the landscape. Populated after calling build_from_data or when loading a graph via build_from_graph if the file contains configs data.

config_dict : dict or None

A dictionary describing the encoding scheme for configuration variables. Keys are typically integer indices of variables, and values are dictionaries specifying properties like 'type' (e.g., 'boolean', 'categorical') and 'max' (maximum encoded value). Populated after calling build_from_data.

data_types : dict or None

A dictionary specifying the data type for each variable in the configuration space (e.g., {'var_0': 'boolean', 'var_1': 'categorical'}). Validated and stored during build_from_data. Required for certain distance calculations.

n_configs : int or None

The total number of configurations (nodes) in the landscape graph. Populated after calling build_from_data or build_from_graph.

n_vars : int or None

The number of variables (dimensions) defining a configuration in the genotype space. Populated after calling build_from_data or inferred by build_from_graph.

n_edges : int or None

The total number of directed edges (connections) in the landscape graph. Populated after calling build_from_data or build_from_graph.

n_lo : int or None

The number of distinct local optima (plateau-aware): each neutral plateau-LO counts once, plus every single-point LO. This is the "number of local optima". Populated after graph analysis.

n_lo_members : int or None

The total number of local-optimum member nodes (every member of every plateau-LO plus single-point LOs); equals len(lo_index). For a landscape with no neutral plateaus this equals n_lo. Populated after graph analysis.

lo_index : list[int] or None

A sorted list of all node indices that are local optima (plateau-aware); its length is n_lo_members. Populated after graph analysis.

go_index : int or None

The node index of the global optimum (the configuration with the highest or lowest fitness). Populated after graph analysis.

go : dict or None

A dictionary containing the attributes (including 'fitness') of the global optimum node. Populated after graph analysis.

lon : ig.graph or None

The Local Optima Network (LON) graph, if calculated via get_lon. Nodes in the LON are local optima from the main landscape graph, and edges represent accessibility between their basins (Ochoa 2021).

has_lon : bool

Flag indicating whether the LON has been calculated and stored in the lon attribute.

maximize : bool

Indicates whether the objective is to maximize (True) or minimize (False) the fitness values. Set during build_from_data or build_from_graph.

epsilon : float

Neutrality threshold. Neighbors with |fitness_a - fitness_b| <= epsilon are classified as neutral rather than improving/worsening. When epsilon > 0, a plateau layer is constructed and downstream analyses become plateau-aware. Defaults to 0.

plateaus : dict or None

Mapping from 0-based plateau ID to the list of member node indices. Only multi-member plateaus are stored; singleton nodes have plateau_id = -1. Populated when epsilon > 0 and neutral pairs exist.

n_plateau : int or None

The number of neutral plateaus (multi-member connected components of neutral neighbors). Populated when epsilon > 0 and neutral pairs exist; 0 when no plateaus; None before construction.

plateau_lo_index : list[int] or None

Plateau IDs (0-based) whose plateaus are local optima. Does not include single-point LOs.

n_plateau_lo : int or None

len(plateau_lo_index).

verbose : bool

The verbosity level set during initialization or construction.

_is_built : bool

Internal flag indicating if the landscape has been populated via build_from_data or build_from_graph.

n_configs

Number of configurations (nodes). Derived from the graph once built (single source of truth); falls back to the working count during construction, before self.graph exists.

n_edges

Number of directed edges. Derived from the graph once built (single source of truth); falls back to the working count during construction.

shape

Return the shape (n_configs, n_edges) of the landscape graph.

configs : Optional[pd.Series]

Per-node configuration tuple Series (built lazily, then cached).

The numeric _configs_array is the source of truth and is what graph construction consumes; this tuple Series is a downstream artifact used only by analyses that need per-node configuration tuples. It is materialised from _configs_array on first access -- builds that never read configs (the common construction-only path) skip the cost entirely. Returns None only when neither the cached Series nor a numeric array is available (e.g. an unbuilt landscape, or a graph load from which configurations could not be reconstructed).

Note: a landscape.configs is None check materialises the Series if _configs_array is present.

basins : pd.Series

Per-node greedy basin size (size_basin_greedy), computed lazily.

accessible_paths : pd.Series

Per-node accessible-basin size (size_basin_accessible), computed lazily.

dist_to_go : pd.Series

Per-node configuration distance to the nearest global optimum (dist_go), computed lazily.

neighbor_fitness : pd.Series

Per-node mean neighbour fitness (mean_neighbor_fit), computed lazily.

pagerank : pd.Series

Per-node PageRank centrality, computed lazily.

PageRank is not used by any landscape metric; it is only an optional, descriptive node attribute. To keep it off the construction critical path (where it was the single dominant cost) it is computed on first access here -- producing values identical to the eager computation (weighted by delta_fit when present, directed=True).

get_params()
get_params() -> Dict[str, Any]

Return the constructor parameters as a dict.

Useful for introspection and reconstruction: type(ls)(**ls.get_params()) yields a fresh, unbuilt landscape with the same configuration.

register_input_handler()
register_input_handler(
    data_type: str, handler: InputHandler
) -> None

Register a custom input handler for a data type on this instance.

register_neighbor_generator()
register_neighbor_generator(
    data_type: str, generator: NeighborGenerator
) -> None

Register a custom neighbor generator for a data type on this instance.

build_from_data()
build_from_data(
    X: Any,
    f: Union[Series, list, ndarray],
    data_types: Optional[Dict[str, str]] = None,
    epsilon: float = 0,
    tau: Optional[float] = None,
    filter_mode: str = "any",
    n_edit: int = 1,
    neighborhood_strategy: Optional[str] = None,
    verbose: Optional[bool] = True,
) -> "Landscape"

Construct the landscape graph and properties from configuration data.

Parameters

X : pandas.DataFrame or numpy.ndarray or a list of strings

The configuration data, where each row represents a genotype or configuration, and columns represent variables or sites. Column names are converted to strings and must remain unique. Names used for landscape attributes, such as 'fitness', 'out_degree' and 'is_lo', are reserved.

f : pandas.Series, list, or numpy.ndarray

The fitness values corresponding to each configuration in X. Must have the same length as X. Values are normalized to float64 before fitness arithmetic.

data_types : dict[str, str], default=None

A dictionary specifying the type of each variable (column in X if DataFrame, or inferred column index if ndarray). Keys must match column names/indices, and values must be one of 'boolean', 'categorical', or 'ordinal'. This information is crucial for determining neighborhood relationships and calculating distances. Optional: the typed landscape subclasses (BooleanLandscape, OrdinalLandscape, DNALandscape/RNALandscape/ ProteinLandscape) auto-detect it. Supply it explicitly for a generic Landscape(kind="default") with heterogeneous (mixed) columns.

epsilon : float, default=0

Neutrality threshold. Neighboring configurations whose absolute fitness difference is <= epsilon are treated as neutral (equal-fitness) rather than strictly improving/worsening. When epsilon > 0, a plateau layer is constructed: neutral neighbors are grouped into plateaus via connected components, and downstream analyses (local optima, basins, accessible paths) become plateau-aware. A higher epsilon produces a smoother landscape with fewer, larger local optima; epsilon=0 preserves strict-inequality behavior (the default).

tau : float, default=None

Functional threshold. Configurations whose fitness is above (when maximize=True) or below (when maximize=False) this value are considered "functional". Used together with filter_mode to focus the landscape on biologically or practically relevant regions.

filter_mode : str, default='any'

How to apply the functional threshold tau. Options:

  • 'any': Pre-construction filter. Remove non-functional configurations before building the graph. If maximize: keep fitness >= tau; if minimize: keep fitness <= tau.

  • 'both': Post-construction filter. Keep all configurations, but remove edges where both endpoints are non-functional (both < tau when maximize, both > tau when minimize). Preserves transitions between functional and non-functional regions.

n_edit : int, default=1

The edit distance defining the neighborhood. For pairwise and broadcast strategies, an undirected neighbor pair is kept when Hamming distance is positive and <= n_edit. The active strategy uses type-specific generators, which only support n_edit=1 for boolean, sequence, and default landscapes (use pairwise or broadcast for multi-edit Hamming graphs).

neighborhood_strategy : str, default=None

Strategy for identifying neighboring configurations. When None (the default), the class default is used (_default_neighborhood_strategy — 'auto' for most landscapes, 'active' for ordinal). On ordinal (or mixed landscapes containing an ordinal variable), 'auto' resolves to 'active' and the Hamming-based 'pairwise'/'broadcast' strategies emit a warning, because they ignore the ±1-step ordinal adjacency. Options:

  • 'auto': Automatically selects the fastest strategy based on dataset size, sequence length, and alphabet size. Uses 'pairwise' when the full distance matrix fits in memory (~4 GiB), falls back to 'broadcast' when vectorized per-config distances are cheaper than candidate generation, and defaults to 'active' otherwise.

  • 'active': For each configuration, enumerates all possible single-edit (or n_edit-edit) mutant neighbors and checks whether they exist in the dataset via hash lookup. Efficient for dense datasets where most proposed neighbors are present.

  • 'pairwise': Computes the full pairwise Hamming distance matrix using scipy.spatial.distance.pdist. Very fast for small-to-moderate datasets (up to ~25 000 configurations) thanks to highly optimized C code. Memory usage scales as O(n_configs^2).

  • 'broadcast': For each configuration, computes its Hamming distance to all other configurations using vectorized NumPy operations. Suitable for large datasets with long sequences and sparse sampling, where the 'active' strategy would waste time generating candidates that don't exist and 'pairwise' would exceed memory limits.

verbose : bool, default=True

If provided, overrides the instance's verbosity setting for this method call. Controls printed output during the build process.

Returns

Landscape

The populated instance itself (self), to support method chaining (e.g. BooleanLandscape().build_from_data(X, f)).

This method takes genotype-phenotype data (configurations X and their corresponding fitness values f) and builds the underlying graph structure of the fitness landscape. Nodes represent configurations, and edges connect neighbors based on the specified edit distance (n_edit). It then determines the basic landscape properties (local optima and the global optimum). More expensive analyses — basins of attraction, accessible paths, distance-to-optimum and neighbour fitness — are computed lazily on first access via the .basins / .accessible_paths / .dist_to_go / .neighbor_fitness properties.

This method populates the core attributes of the Landscape instance.

Notes

The construction process involves several steps:

  1. Resolve the type-specific preparation and neighbor-generation strategies.

  2. Apply the optional pre-construction fitness filter.

  3. Prepare the raw input based on the landscape type.

  4. Warn on missing values (no automatic row drops); drop duplicate configurations; encode step errors if NaNs remain.

  5. Encode the data and cache configuration metadata.

  6. Construct the core directed landscape graph.

  7. Apply post-construction pruning and remap cached metadata.

  8. Build neutral plateaus and analyze derived landscape properties.

Raises

RuntimeError

If the build_from_data or build_from_graph method has already been called on this instance. Create a new instance to rebuild.

ValueError

If input data X, f, or data_types are invalid (e.g., mismatched lengths, empty data, invalid types in data_types), or if the construction process fails (e.g., resulting graph is empty).

TypeError

If input types are incorrect (e.g., data_types is not a dict).

get_data()
get_data(
    lo_only: bool = False, include_pagerank: bool = False
) -> pd.DataFrame

Extracts landscape data as a pandas DataFrame.

Parameters

lo_only : bool, default=False

If True, returns data only for the configurations identified as local optima. If the Local Optima Network (LON) has been computed (see get_lon), data from the LON graph is returned. Otherwise, it returns data from the main graph filtered for local optima nodes. If False, returns data for all configurations in the main graph.

include_pagerank : bool, default=False

If True, compute (if needed) and include a pagerank column. PageRank is otherwise omitted -- it is an optional, comparatively expensive centrality that most callers do not need, so get_data no longer triggers it as a hidden side effect.

Returns

pandas.DataFrame

A DataFrame containing the attributes of the landscape nodes. Index corresponds to the node indices.

Returns a DataFrame where rows correspond to configurations (nodes) and columns correspond to their attributes (e.g., fitness, degree, basin information, original features). Feature columns appear in the original input order.

Raises

RuntimeError

If the landscape has not been built (via build_from_data or build_from_graph) before calling this method. If lo_only=True and the LON graph (self.lon) is unexpectedly None despite self.has_lon being True.

get_lon()
get_lon(
    mlon: bool = True,
    min_edge_freq: int = 3,
    trim: Optional[int] = None,
    verbose: Optional[bool] = None,
) -> ig.Graph

Constructs and returns the Local Optima Network (LON).

Parameters

mlon : bool, default=True

If True, also build the monotonic LON (edges restricted to non-worsening transitions).

min_edge_freq : int, default=3

Keep a LON edge only when the number of basin transitions between two optima is strictly greater than this threshold.

trim : int, default=None

If given, keep only the trim strongest outgoing edges per node.

verbose : bool, default=None

Verbosity override; defaults to the landscape's own verbose.

Returns

ig.Graph

The constructed Local Optima Network graph. The graph is also stored in the self.lon attribute, and self.has_lon is set to True.

The LON is a coarse-grained representation of the fitness landscape where nodes are the local optima of the original landscape, and edges represent the possibility of transitions between their basins of attraction, typically weighted by the fitness difference or distance between the optima. This method requires the landscape graph to be built and local optima to be identified.

The landscape's own graph, configurations, optima and config_dict are supplied automatically; the parameters below control the LON itself and are forwarded to graphfla.lon.get_lon.

Raises

RuntimeError

If the landscape has not been built, or if essential attributes (graph, configs, lo_index, config_dict) required for LON construction are missing.

describe()
describe() -> Dict[str, Any]

Return a structured summary of the landscape as a dict.

Returns

dict

class, kind, built, maximize and epsilon are always present. When built, size/optima fields (n_vars, n_configs, n_edges, n_lo, go_index), the calculation flags, and -- if a plateau layer exists -- the plateau counts are added.

Unlike a printed summary, the returned mapping is composable and testable -- callers can log it, assert on it, or render it. For a quick human-readable view use print(landscape) (see __str__).

build_from_graph()

Inherited from _IOMixin

build_from_graph(
    filepath: str, *, verbose: bool = True
) -> "Landscape"

Construct a landscape from a saved graph file.

Parameters

filepath : str

Path to the saved graph file (.graphml).

verbose : bool, default=True

Controls verbosity of output during loading and analysis.

Returns

Landscape

A new instance populated with the graph and inferred properties.

This class method creates a new landscape instance by loading a previously saved graph, avoiding the need to reconstruct the landscape from original configuration data. This is significantly faster than building from scratch.

Notes

This method will:

  1. Load the saved graph structure and attributes

  2. Infer essential landscape properties from the graph

  3. Recalculate local optima and global optimum from the graph structure

Previously computed attributes (basins, accessible paths, distances, neighbor fitness) are preserved from the saved graph if present.

Some specialized attributes from subclasses (like sequence_length in SequenceLandscape) will be inferred where possible.

Only load GraphML from trusted sources: embedded configuration metadata is parsed with ast.literal_eval (safe against arbitrary code execution, but not a substitute for validating untrusted files).

Raises

ValueError

If the file cannot be read or doesn't contain valid graph data.

FileNotFoundError

If the specified file doesn't exist.

to_graph()

Inherited from _IOMixin

to_graph(filepath: str) -> None

Save the landscape graph and essential attributes to a file.

Parameters

filepath : str

The path where the graph file will be saved. If the file doesn't end with '.graphml', this extension will be added automatically.

This method serializes the landscape's graph structure and relevant attributes to a GraphML file, which can later be loaded using build_from_graph. This allows efficient storage and sharing of landscapes without requiring re-construction from scratch.

Notes

The GraphML format preserves the graph structure and all vertex/edge attributes. In addition to the graph itself, essential landscape attributes like maximize and epsilon are stored as graph attributes.

Raises

NotBuiltError

If the landscape has not been built.

ValueError

If the graph cannot be saved to the specified path.

References

[Wright 1932] Wright, S. The roles of mutation, inbreeding, crossbreeding and selection in evolution. Proceedings Sixth International Congress Genetics 1, 356-366 (1932).

[Papkou 2023] Papkou, A. et al. A rugged yet easily navigable fitness landscape. Science 382, eadh3860 (2023).

[Li 2016] Li, C. et al. The fitness landscape of a tRNA gene. Science 352, 837-840 (2016).

[Puchta 2016] Puchta, O. et al. Network of epistatic interactions within a yeast snoRNA. Science 352, 840-844 (2016).

[Poelwijk 2007] Poelwijk, F. J. et al. Empirical fitness landscapes reveal accessible evolutionary paths. Nature 445, 383-386 (2007).

[Carneiro 2010] Carneiro, M. & Hartl, D. L. Adaptive landscapes and protein evolution. PNAS 107, 1747-1751 (2010).

[Ochoa 2021] Ochoa, G. et al. Local optima networks: A survey. Journal of Heuristics 27, 79-134 (2021).

Examples

>>> import pandas as pd
>>> import numpy as np
>>> X_data = pd.DataFrame({'var_0': [0, 0, 1, 1], 'var_1': [0, 1, 0, 1]})
>>> f_data = pd.Series([1.0, 2.0, 3.0, 2.5])
>>> landscape = BooleanLandscape().build_from_data(X_data, f_data, verbose=False)
>>> landscape.describe()["n_lo"]
1
>>> repr(landscape)
'BooleanLandscape(maximize=True)'
>>> print(f"Number of configurations: {landscape.n_configs}")
Number of configurations: 4
>>> print(f"Global optimum fitness: {landscape.go['fitness']}")  
Global optimum fitness: 3.0