Skip to content

Materials design

Hybrid perovskites · ABX₃

In this tutorial, we’ll use GraphFLA to explore how a material’s ingredients affect its electronic properties. To work through a materials design problem, we’ll build a landscape of hybrid perovskites and examine how constituent choices influence the band gap.

Open In Colab

Download notebook Notebook + data

Run the notebook in Google Colab, or download it to run locally. It downloads its data when no data/ folder is beside it.

1. Setting up

On Google Colab, the cell below installs GraphFLA. In any environment, it also downloads the dataset into a data/ folder next to the notebook if the files are not already there.

We then import GraphFLA for landscape construction and analysis, and pandas for working with the data.

import sys
from pathlib import Path
from urllib.request import urlretrieve

if "google.colab" in sys.modules:
    %pip install -q graphfla==0.4.0

DATA_URL = "https://raw.githubusercontent.com/COLA-Laboratory/GraphFLA/v0.4.0/tutorials/datasets/data/"
Path("data").mkdir(exist_ok=True)
for name in ["perovskites.csv"]:
    if not Path("data", name).exists():
        urlretrieve(DATA_URL + name, Path("data", name))
from graphfla import analysis
from graphfla.landscape import Landscape
import pandas as pd

2. Loading the dataset

Hybrid organic-inorganic perovskites combine organic ions with inorganic constituents in a crystal. Changing these ingredients changes their electronic properties. To explore this design space, we’ll use a benchmark derived from the calculations of Kim et al. (2017). It contains 192 ABX₃ formulas, each with an organic A-site cation, a B-site metal and a halide.

We’ll minimise the calculated electronic band gap, the energy gap between valence and conduction bands. The benchmark uses the HSE06 calculation method; a smaller gap alone does not establish better solar-cell performance.

Column Meaning
organic One of 16 organic A-site cations.
metal Ge, Sn or Pb at the B site.
halide F, Cl, Br or I.
hse_gap_eV HSE06 band gap in eV.

Let’s load the Olympus lookup. The original study included multiple structures per formula; the lookup’s rule for selecting one value per formula is undocumented.

df = pd.read_csv(
    "data/perovskites.csv", header=None,
    names=["organic", "metal", "halide", "hse_gap_eV"],
)
df.head()

Output

organic metal halide hse_gap_eV
0 ethylammonium Ge F 5.3704
1 ethylammonium Ge Cl 3.1393
2 ethylammonium Ge Br 2.7138
3 ethylammonium Ge I 2.2338
4 ethylammonium Sn F 3.9789

3. Preparing the inputs

To construct a landscape, GraphFLA needs two aligned inputs:

  • X: one row per configuration and one column per variable.
  • f: one measured or calculated outcome for each row of X.

The three constituent identities form X, and the band gap forms f. Stoichiometric roles stay fixed at ABX₃. We are selecting ingredients rather than adjusting metal fractions in an alloy.

X = df[["organic", "metal", "halide"]]
f = df["hse_gap_eV"]
X.head()

Output

organic metal halide
0 ethylammonium Ge F
1 ethylammonium Ge Cl
2 ethylammonium Ge Br
3 ethylammonium Ge I
4 ethylammonium Sn F

4. Constructing the landscape

We can use Landscape with three categorical variables. Neighbours replace one constituent while leaving the other two identities fixed. Each replacement counts as one step, regardless of chemical similarity.

We set maximize=False because the benchmark seeks a smaller band gap. Improving edges point towards lower gaps, while stored values remain in eV.

landscape = Landscape(maximize=False)
landscape.build_from_data(
    X, f, data_types={column: "categorical" for column in X},
    neighborhood_strategy="active", verbose=False,
)

Output

Landscape(kind='default', maximize=False)

5. Inspecting the landscape

Let’s first look at what we’ve built. Printing the landscape shows its variables, configurations, improving edges and local optima:

print(landscape)

Output

Landscape(kind='default'): 3 variables, 192 configurations, 1920 edges, 1 local optima

For a closer look at individual perovskite formulas, we can call get_data(). The outcome appears as fitness; out_degree counts directly improving moves, and is_lo indicates local-optimum membership.

landscape.get_data().head()

Output

organic metal halide fitness in_degree out_degree is_lo
0 ethylammonium Ge F 5.3704 7 13 False
1 ethylammonium Ge Cl 3.1393 12 8 False
2 ethylammonium Ge Br 2.7138 13 7 False
3 ethylammonium Ge I 2.2338 11 9 False
4 ethylammonium Sn F 3.9789 15 5 False

We can also summarise the calculated band gaps with pandas:

landscape.get_data()["fitness"].describe()

Output

count    192.000000
mean       3.332590
std        1.139340
min        1.524900
25%        2.456575
50%        3.079100
75%        4.031850
max        6.324200
Name: fitness, dtype: float64

To inspect the formula with the lowest recorded band gap, we can use the global-optimum node index:

landscape[landscape.go_index]

Output

{'organic': 'hydrazinium',
 'metal': 'Sn',
 'halide': 'I',
 'fitness': np.float64(1.5249),
 'in_degree': 20,
 'out_degree': 0,
 'is_lo': True}

6. Analysing the landscape

Next, we’ll count local optima, examine local fitness similarity and interaction orders, and assess the trend towards the best observed configuration.

6.1 Number of local optima

We can start with landscape.n_lo. Each connected neutral optimum plateau counts once. A local optimum has no improving move out of it, including through its neutral plateau. Multiple optima mean successive improvements can end at different configurations.

print(f"Number of local optima: {landscape.n_lo}")

Output

Number of local optima: 1

6.2 Fitness autocorrelation

How similar are responses at neighbouring steps? autocorrelation() measures persistence along sampled walks. Higher values indicate more persistence under the same walk settings.

We use 200 walks of up to 20 visited configurations, lag one and a fixed seed. Walks traverse stored edges in either direction; separately stored neutral pairs are excluded.

walk_autocorrelation = analysis.autocorrelation(
    landscape, walk_length=20, walk_times=200, lag=1, seed=42,
)
print(f"Lag-1 fitness autocorrelation: {walk_autocorrelation:.4f}")

Output

Lag-1 fitness autocorrelation: 0.6915

6.3 Walsh–Hadamard decomposition

A metal substitution may affect the gap differently with different halides. walsh_hadamard() fits individual-constituent effects and interactions between pairs. Compare first- and second-order training r2; delta_r2 shows the additional variation explained by pairwise terms.

Three-way effects remain outside the fit. Coefficients describe the original band-gap scale even though we minimise it. The model’s uniform-product variance spectrum is distinct from the incremental training fit and does not establish predictive accuracy.

wh = analysis.walsh_hadamard(landscape, max_order=2)
wh["order_summary"]

Output

order r2 delta_r2 rmse n_terms rank alpha model_variance_fraction n_nonzero
0 0 0.000000 0.000000 1.136369 1 1 NaN 0.000000 <NA>
1 1 0.943964 0.943964 0.269000 21 21 NaN 0.958162 <NA>
2 2 0.985182 0.041218 0.138328 102 102 NaN 0.041838 <NA>

We can also inspect the largest fitted nonconstant coefficients. Their units follow the response, and their signs depend on the encoded contrasts:

wh["coefficients"].query("order > 0").sort_values(
    "coefficient", key=abs, ascending=False,
).head(8)

Output

order positions term coefficient
3 1 (3,) F_3_I -2.745577
1 1 (3,) F_3_Br -2.213302
2 1 (3,) F_3_Cl -1.610529
70 2 (1, 2) ethylammonium_1_hydroxylammonium-Ge_2_Pb 1.246250
94 2 (1, 3) ethylammonium_1_tetramethylammonium-F_3_I -1.176033
65 2 (1, 2) ethylammonium_1_hydrazinium-Ge_2_Pb 1.071975
35 2 (1, 2) ethylammonium_1_ammonium-Ge_2_Pb 1.007050
19 1 (1,) ethylammonium_1_tetramethylammonium 0.974258

6.4 Fitness-distance correlation

Distance counts how many constituent identities differ from the selected best formula. It does not encode structural or electronic similarity.

We use Spearman fdc to compare ranks. A positive value means lower responses tend to occur nearer the optimum. A value near zero indicates little monotonic association. This trend does not guarantee an improving path from every configuration.

fitness_distance_r = analysis.fdc(landscape, method="spearman")
print(f"Fitness-distance correlation: {fitness_distance_r:.4f}")

Output

Fitness-distance correlation: 0.5994

This graph has one local optimum. The additive fit already explains about 94.4% of response variation, increasing to 98.5% with pairwise terms. Positive FDC (0.599) matches the minimisation objective.

7. Running several analyses together

We can collect autocorrelation and FDC with analysis.profile(), keeping the same walk settings. landscape.n_lo supplies the count directly; the W–H result keeps its separate coefficient and order-summary tables.

analysis.profile(
    landscape,
    metrics=["autocorrelation", "fdc"],
    params={"autocorrelation": {"walk_length": 20, "walk_times": 200, "lag": 1}},
    seed=42, progress=False,
)

Output

autocorrelation    0.691489
fdc                0.599377
dtype: float64

8. Analysis reference

For further exploration, analysis.list_metrics() lists the metrics available to profile(). The worked W–H result also exposes coefficients and fit_info; its order summary separates cumulative training fit from the fitted model’s variance spectrum.

Data sources

Kim C. et al. (2017). A hybrid organic-inorganic perovskite dataset. This tutorial uses the Olympus perovskites lookup, containing one HSE06 gap per formula, rather than all 1,346 relaxed structures in the original study.