OrdinalLandscape ¶
OrdinalLandscape(maximize: bool = True)A specialized landscape class for ordinal configuration spaces.
Each configuration is a vector of ordered discrete values — for
example, dose levels {0, 1, 2, 3}, hyperparameter tiers, Likert
scales, or any discrete-but-ordered factor.
Unlike a categorical landscape, the neighborhood of an ordinal configuration is ±1 step on the ordinal scale at one position — i.e., Manhattan-distance-1 on each axis — and distances to the global optimum are computed using Manhattan rather than Hamming distance. This matches the standard definition used in the ordinal-landscape literature.
The constructor takes no shape arguments: the number of levels for
each variable is auto-detected from the data during
build_from_data (it follows pandas.Categorical(...).codes,
so the natural order of integer-coded values is preserved).
Methods at a glance
build_from_graph()- Construct a landscape from a saved graph file.
to_graph()- Save the landscape graph and essential attributes to a file.
get_params()- Return the constructor parameters as a dict.
register_input_handler()- Register a custom input handler for a data type on this instance.
register_neighbor_generator()- Register a custom neighbor generator for a data type on this instance.
build_from_data()- Construct the landscape graph and properties from configuration data.
get_data()- Extracts landscape data as a pandas DataFrame.
get_lon()- Constructs and returns the Local Optima Network (LON).
describe()- Return a structured summary of the landscape as a dict.
Parameters
-
maximize: bool, default=True Determines the optimization direction. If True, the landscape seeks higher fitness values. If False, it seeks lower values.
Parameters
-
maximize: bool, default=True Determines the optimization direction.
n_configs
¶
Number of configurations (nodes). Derived from the graph once built
(single source of truth); falls back to the working count during
construction, before self.graph exists.
n_edges
¶
Number of directed edges. Derived from the graph once built (single source of truth); falls back to the working count during construction.
shape
¶
Return the shape (n_configs, n_edges) of the landscape graph.
configs
: Optional[pd.Series]
¶
Per-node configuration tuple Series (built lazily, then cached).
The numeric _configs_array is the source of truth and is what graph
construction consumes; this tuple Series is a downstream artifact
used only by analyses that need per-node configuration tuples. It is
materialised from _configs_array on first access -- builds that never
read configs (the common construction-only path) skip the cost
entirely. Returns None only when neither the cached Series nor a
numeric array is available (e.g. an unbuilt landscape, or a graph load
from which configurations could not be reconstructed).
Note: a landscape.configs is None check materialises the Series
if _configs_array is present.
basins
: pd.Series
¶
Per-node greedy basin size (size_basin_greedy), computed lazily.
accessible_paths
: pd.Series
¶
Per-node accessible-basin size (size_basin_accessible), computed lazily.
dist_to_go
: pd.Series
¶
Per-node configuration distance to the nearest global optimum
(dist_go), computed lazily.
neighbor_fitness
: pd.Series
¶
Per-node mean neighbour fitness (mean_neighbor_fit), computed lazily.
pagerank
: pd.Series
¶
Per-node PageRank centrality, computed lazily.
PageRank is not used by any landscape metric; it is only an optional,
descriptive node attribute. To keep it off the construction critical
path (where it was the single dominant cost) it is computed on first
access here -- producing values identical to the eager computation
(weighted by delta_fit when present, directed=True).
build_from_graph()
build_from_graph(
filepath: str, *, verbose: bool = True
) -> "Landscape"Construct a landscape from a saved graph file.
Parameters
-
filepath: str Path to the saved graph file (.graphml).
-
verbose: bool, default=True Controls verbosity of output during loading and analysis.
Returns
-
Landscape A new instance populated with the graph and inferred properties.
This class method creates a new landscape instance by loading a previously saved graph, avoiding the need to reconstruct the landscape from original configuration data. This is significantly faster than building from scratch.
Notes
This method will:
-
Load the saved graph structure and attributes
-
Infer essential landscape properties from the graph
-
Recalculate local optima and global optimum from the graph structure
Previously computed attributes (basins, accessible paths, distances, neighbor fitness) are preserved from the saved graph if present.
Some specialized attributes from subclasses (like sequence_length in SequenceLandscape) will be inferred where possible.
Only load GraphML from trusted sources: embedded configuration metadata
is parsed with ast.literal_eval (safe against arbitrary code
execution, but not a substitute for validating untrusted files).
Raises
-
ValueError If the file cannot be read or doesn't contain valid graph data.
-
FileNotFoundError If the specified file doesn't exist.
to_graph()
to_graph(filepath: str) -> NoneSave the landscape graph and essential attributes to a file.
Parameters
-
filepath: str The path where the graph file will be saved. If the file doesn't end with '.graphml', this extension will be added automatically.
This method serializes the landscape's graph structure and relevant
attributes to a GraphML file, which can later be loaded using build_from_graph.
This allows efficient storage and sharing of landscapes without requiring
re-construction from scratch.
Notes
The GraphML format preserves the graph structure and all vertex/edge attributes.
In addition to the graph itself, essential landscape attributes like maximize
and epsilon are stored as graph attributes.
Raises
-
NotBuiltError If the landscape has not been built.
-
ValueError If the graph cannot be saved to the specified path.
build_from_data()
build_from_data(
X: Any,
f: Union[Series, list, ndarray],
data_types: Optional[Dict[str, str]] = None,
epsilon: float = 0,
tau: Optional[float] = None,
filter_mode: str = "any",
n_edit: int = 1,
neighborhood_strategy: Optional[str] = None,
verbose: Optional[bool] = True,
) -> "Landscape"Construct the landscape graph and properties from configuration data.
Parameters
-
X: pandas.DataFrame or numpy.ndarray or a list of strings The configuration data, where each row represents a genotype or configuration, and columns represent variables or sites. Column names are converted to strings and must remain unique. Names used for landscape attributes, such as 'fitness', 'out_degree' and 'is_lo', are reserved.
-
f: pandas.Series, list, or numpy.ndarray The fitness values corresponding to each configuration in
X. Must have the same length asX. Values are normalized to float64 before fitness arithmetic.-
data_types: dict[str, str], default=None A dictionary specifying the type of each variable (column in
Xif DataFrame, or inferred column index if ndarray). Keys must match column names/indices, and values must be one of 'boolean', 'categorical', or 'ordinal'. This information is crucial for determining neighborhood relationships and calculating distances. Optional: the typed landscape subclasses (BooleanLandscape,OrdinalLandscape,DNALandscape/RNALandscape/ProteinLandscape) auto-detect it. Supply it explicitly for a genericLandscape(kind="default")with heterogeneous (mixed) columns.-
epsilon: float, default=0 Neutrality threshold. Neighboring configurations whose absolute fitness difference is
<= epsilonare treated as neutral (equal-fitness) rather than strictly improving/worsening. Whenepsilon > 0, a plateau layer is constructed: neutral neighbors are grouped into plateaus via connected components, and downstream analyses (local optima, basins, accessible paths) become plateau-aware. A higher epsilon produces a smoother landscape with fewer, larger local optima;epsilon=0preserves strict-inequality behavior (the default).-
tau: float, default=None Functional threshold. Configurations whose fitness is above (when
maximize=True) or below (whenmaximize=False) this value are considered "functional". Used together withfilter_modeto focus the landscape on biologically or practically relevant regions.-
filter_mode: str, default='any' How to apply the functional threshold
tau. Options:-
'any': Pre-construction filter. Remove non-functional configurations before building the graph. If maximize: keep fitness >= tau; if minimize: keep fitness <= tau. -
'both': Post-construction filter. Keep all configurations, but remove edges where both endpoints are non-functional (both < tau when maximize, both > tau when minimize). Preserves transitions between functional and non-functional regions.
-
-
n_edit: int, default=1 The edit distance defining the neighborhood. For
pairwiseandbroadcaststrategies, an undirected neighbor pair is kept when Hamming distance is positive and<= n_edit. Theactivestrategy uses type-specific generators, which only supportn_edit=1for boolean, sequence, and default landscapes (usepairwiseorbroadcastfor multi-edit Hamming graphs).-
neighborhood_strategy: str, default=None Strategy for identifying neighboring configurations. When
None(the default), the class default is used (_default_neighborhood_strategy—'auto'for most landscapes,'active'for ordinal). On ordinal (or mixed landscapes containing an ordinal variable),'auto'resolves to'active'and the Hamming-based'pairwise'/'broadcast'strategies emit a warning, because they ignore the ±1-step ordinal adjacency. Options:-
'auto': Automatically selects the fastest strategy based on dataset size, sequence length, and alphabet size. Uses'pairwise'when the full distance matrix fits in memory (~4 GiB), falls back to'broadcast'when vectorized per-config distances are cheaper than candidate generation, and defaults to'active'otherwise. -
'active': For each configuration, enumerates all possible single-edit (orn_edit-edit) mutant neighbors and checks whether they exist in the dataset via hash lookup. Efficient for dense datasets where most proposed neighbors are present. -
'pairwise': Computes the full pairwise Hamming distance matrix usingscipy.spatial.distance.pdist. Very fast for small-to-moderate datasets (up to ~25 000 configurations) thanks to highly optimized C code. Memory usage scales asO(n_configs^2). -
'broadcast': For each configuration, computes its Hamming distance to all other configurations using vectorized NumPy operations. Suitable for large datasets with long sequences and sparse sampling, where the'active'strategy would waste time generating candidates that don't exist and'pairwise'would exceed memory limits.
-
-
verbose: bool, default=True If provided, overrides the instance's verbosity setting for this method call. Controls printed output during the build process.
Returns
-
Landscape The populated instance itself (
self), to support method chaining (e.g.BooleanLandscape().build_from_data(X, f)).
This method takes genotype-phenotype data (configurations X and their
corresponding fitness values f) and builds the underlying graph
structure of the fitness landscape. Nodes represent configurations, and
edges connect neighbors based on the specified edit distance (n_edit).
It then determines the basic landscape properties (local optima and the
global optimum). More expensive analyses — basins of attraction,
accessible paths, distance-to-optimum and neighbour fitness — are computed
lazily on first access via the .basins / .accessible_paths /
.dist_to_go / .neighbor_fitness properties.
This method populates the core attributes of the Landscape instance.
Notes
The construction process involves several steps:
-
Resolve the type-specific preparation and neighbor-generation strategies.
-
Apply the optional pre-construction fitness filter.
-
Prepare the raw input based on the landscape type.
-
Warn on missing values (no automatic row drops); drop duplicate configurations; encode step errors if NaNs remain.
-
Encode the data and cache configuration metadata.
-
Construct the core directed landscape graph.
-
Apply post-construction pruning and remap cached metadata.
-
Build neutral plateaus and analyze derived landscape properties.
Raises
-
RuntimeError If the
build_from_dataorbuild_from_graphmethod has already been called on this instance. Create a new instance to rebuild.-
ValueError If input data
X,f, ordata_typesare invalid (e.g., mismatched lengths, empty data, invalid types indata_types), or if the construction process fails (e.g., resulting graph is empty).-
TypeError If input types are incorrect (e.g.,
data_typesis not a dict).
get_data()
get_data(
lo_only: bool = False, include_pagerank: bool = False
) -> pd.DataFrameExtracts landscape data as a pandas DataFrame.
Parameters
-
lo_only: bool, default=False If True, returns data only for the configurations identified as local optima. If the Local Optima Network (LON) has been computed (see
get_lon), data from the LON graph is returned. Otherwise, it returns data from the main graph filtered for local optima nodes. If False, returns data for all configurations in the main graph.-
include_pagerank: bool, default=False If True, compute (if needed) and include a
pagerankcolumn. PageRank is otherwise omitted -- it is an optional, comparatively expensive centrality that most callers do not need, soget_datano longer triggers it as a hidden side effect.
Returns
-
pandas.DataFrame A DataFrame containing the attributes of the landscape nodes. Index corresponds to the node indices.
Returns a DataFrame where rows correspond to configurations (nodes) and columns correspond to their attributes (e.g., fitness, degree, basin information, original features). Feature columns appear in the original input order.
Raises
-
RuntimeError If the landscape has not been built (via
build_from_dataorbuild_from_graph) before calling this method. Iflo_only=Trueand the LON graph (self.lon) is unexpectedly None despiteself.has_lonbeing True.
get_lon()
get_lon(
mlon: bool = True,
min_edge_freq: int = 3,
trim: Optional[int] = None,
verbose: Optional[bool] = None,
) -> ig.GraphConstructs and returns the Local Optima Network (LON).
Parameters
-
mlon: bool, default=True If True, also build the monotonic LON (edges restricted to non-worsening transitions).
-
min_edge_freq: int, default=3 Keep a LON edge only when the number of basin transitions between two optima is strictly greater than this threshold.
-
trim: int, default=None If given, keep only the
trimstrongest outgoing edges per node.-
verbose: bool, default=None Verbosity override; defaults to the landscape's own
verbose.
Returns
-
ig.Graph The constructed Local Optima Network graph. The graph is also stored in the
self.lonattribute, andself.has_lonis set to True.
The LON is a coarse-grained representation of the fitness landscape where nodes are the local optima of the original landscape, and edges represent the possibility of transitions between their basins of attraction, typically weighted by the fitness difference or distance between the optima. This method requires the landscape graph to be built and local optima to be identified.
The landscape's own graph, configurations, optima and config_dict
are supplied automatically; the parameters below control the LON itself
and are forwarded to graphfla.lon.get_lon.
Raises
-
RuntimeError If the landscape has not been built, or if essential attributes (
graph,configs,lo_index,config_dict) required for LON construction are missing.
describe()
describe() -> Dict[str, Any]Return a structured summary of the landscape as a dict.
Returns
-
dict class,kind,built,maximizeandepsilonare always present. When built, size/optima fields (n_vars,n_configs,n_edges,n_lo,go_index), the calculation flags, and -- if a plateau layer exists -- the plateau counts are added.
Unlike a printed summary, the returned mapping is composable and
testable -- callers can log it, assert on it, or render it. For a quick
human-readable view use print(landscape) (see __str__).
Notes
-
The
"active"neighborhood strategy is the default for ordinal landscapes (via_default_neighborhood_strategy); it calls theOrdinalNeighborGenerator, which enforces the correct ±1 step (Manhattan-1) semantics. The Hamming-based"pairwise"/"broadcast"strategies would instead treat any single-position change as adjacent — including pairs many steps apart on the ordinal scale — sobuild_from_datawarns if you request them on an ordinal landscape. -
If your data uses non-integer ordered values (e.g., strings such as
"low","mid","high"), pandas will fall back to lexicographic sorting, which is almost never what you want. Either pre-encode the values as integers or pass each column as apandas.Categorical(col, ordered=True, categories=[...])with the explicit order. -
For a mixed landscape combining ordinal variables with boolean or categorical ones, use the generic
Landscapewithkind="default"and supply an explicitdata_typesdictionary.
Examples
>>> import pandas as pd
>>> X = pd.DataFrame({
... "dose_A": [0, 0, 1, 1, 2, 2],
... "dose_B": [0, 1, 0, 1, 0, 1],
... })
>>> f = pd.Series([0.1, 0.3, 0.5, 0.4, 0.7, 0.6])
>>> landscape = OrdinalLandscape(maximize=True).build_from_data(X, f, verbose=False)
Initialize an ordinal landscape.