Materials design¶
Hybrid perovskites · ABX₃
In this tutorial, we’ll use GraphFLA to explore how a material’s ingredients affect its electronic properties. To work through a materials design problem, we’ll build a landscape of hybrid perovskites and examine how constituent choices influence the band gap.
Download notebook Notebook + data
Run the notebook in Google Colab, or download it to run locally. It downloads its data when no data/ folder is beside it.
1. Setting up¶
On Google Colab, the cell below installs GraphFLA. In any environment, it also downloads the dataset into a data/ folder next to the notebook if the files are not already there.
We then import GraphFLA for landscape construction and analysis, and pandas for working with the data.
import sys
from pathlib import Path
from urllib.request import urlretrieve
if "google.colab" in sys.modules:
%pip install -q graphfla==0.4.0
DATA_URL = "https://raw.githubusercontent.com/COLA-Laboratory/GraphFLA/v0.4.0/tutorials/datasets/data/"
Path("data").mkdir(exist_ok=True)
for name in ["perovskites.csv"]:
if not Path("data", name).exists():
urlretrieve(DATA_URL + name, Path("data", name))
from graphfla import analysis
from graphfla.landscape import Landscape
import pandas as pd
2. Loading the dataset¶
Hybrid organic-inorganic perovskites combine organic ions with inorganic constituents in a crystal. Changing these ingredients changes their electronic properties. To explore this design space, we’ll use a benchmark derived from the calculations of Kim et al. (2017). It contains 192 ABX₃ formulas, each with an organic A-site cation, a B-site metal and a halide.
We’ll minimise the calculated electronic band gap, the energy gap between valence and conduction bands. The benchmark uses the HSE06 calculation method; a smaller gap alone does not establish better solar-cell performance.
| Column | Meaning |
|---|---|
organic |
One of 16 organic A-site cations. |
metal |
Ge, Sn or Pb at the B site. |
halide |
F, Cl, Br or I. |
hse_gap_eV |
HSE06 band gap in eV. |
Let’s load the Olympus lookup. The original study included multiple structures per formula; the lookup’s rule for selecting one value per formula is undocumented.
df = pd.read_csv(
"data/perovskites.csv", header=None,
names=["organic", "metal", "halide", "hse_gap_eV"],
)
df.head()
Output
| organic | metal | halide | hse_gap_eV | |
|---|---|---|---|---|
| 0 | ethylammonium | Ge | F | 5.3704 |
| 1 | ethylammonium | Ge | Cl | 3.1393 |
| 2 | ethylammonium | Ge | Br | 2.7138 |
| 3 | ethylammonium | Ge | I | 2.2338 |
| 4 | ethylammonium | Sn | F | 3.9789 |
3. Preparing the inputs¶
To construct a landscape, GraphFLA needs two aligned inputs:
X: one row per configuration and one column per variable.f: one measured or calculated outcome for each row ofX.
The three constituent identities form X, and the band gap forms f. Stoichiometric roles stay fixed at ABX₃. We are selecting ingredients rather than adjusting metal fractions in an alloy.
X = df[["organic", "metal", "halide"]]
f = df["hse_gap_eV"]
X.head()
Output
| organic | metal | halide | |
|---|---|---|---|
| 0 | ethylammonium | Ge | F |
| 1 | ethylammonium | Ge | Cl |
| 2 | ethylammonium | Ge | Br |
| 3 | ethylammonium | Ge | I |
| 4 | ethylammonium | Sn | F |
4. Constructing the landscape¶
We can use Landscape with three categorical variables. Neighbours replace one constituent while leaving the other two identities fixed. Each replacement counts as one step, regardless of chemical similarity.
We set maximize=False because the benchmark seeks a smaller band gap. Improving edges point towards lower gaps, while stored values remain in eV.
landscape = Landscape(maximize=False)
landscape.build_from_data(
X, f, data_types={column: "categorical" for column in X},
neighborhood_strategy="active", verbose=False,
)
Output
Landscape(kind='default', maximize=False)
5. Inspecting the landscape¶
Let’s first look at what we’ve built. Printing the landscape shows its variables, configurations, improving edges and local optima:
print(landscape)
Output
Landscape(kind='default'): 3 variables, 192 configurations, 1920 edges, 1 local optima
For a closer look at individual perovskite formulas, we can call get_data(). The outcome appears as fitness; out_degree counts directly improving moves, and is_lo indicates local-optimum membership.
landscape.get_data().head()
Output
| organic | metal | halide | fitness | in_degree | out_degree | is_lo | |
|---|---|---|---|---|---|---|---|
| 0 | ethylammonium | Ge | F | 5.3704 | 7 | 13 | False |
| 1 | ethylammonium | Ge | Cl | 3.1393 | 12 | 8 | False |
| 2 | ethylammonium | Ge | Br | 2.7138 | 13 | 7 | False |
| 3 | ethylammonium | Ge | I | 2.2338 | 11 | 9 | False |
| 4 | ethylammonium | Sn | F | 3.9789 | 15 | 5 | False |
We can also summarise the calculated band gaps with pandas:
landscape.get_data()["fitness"].describe()
Output
count 192.000000
mean 3.332590
std 1.139340
min 1.524900
25% 2.456575
50% 3.079100
75% 4.031850
max 6.324200
Name: fitness, dtype: float64
To inspect the formula with the lowest recorded band gap, we can use the global-optimum node index:
landscape[landscape.go_index]
Output
{'organic': 'hydrazinium',
'metal': 'Sn',
'halide': 'I',
'fitness': np.float64(1.5249),
'in_degree': 20,
'out_degree': 0,
'is_lo': True}
6. Analysing the landscape¶
Next, we’ll count local optima, examine local fitness similarity and interaction orders, and assess the trend towards the best observed configuration.
6.1 Number of local optima¶
We can start with landscape.n_lo. Each connected neutral optimum plateau counts once. A local optimum has no improving move out of it, including through its neutral plateau. Multiple optima mean successive improvements can end at different configurations.
print(f"Number of local optima: {landscape.n_lo}")
Output
Number of local optima: 1
6.2 Fitness autocorrelation¶
How similar are responses at neighbouring steps? autocorrelation() measures persistence along sampled walks. Higher values indicate more persistence under the same walk settings.
We use 200 walks of up to 20 visited configurations, lag one and a fixed seed. Walks traverse stored edges in either direction; separately stored neutral pairs are excluded.
walk_autocorrelation = analysis.autocorrelation(
landscape, walk_length=20, walk_times=200, lag=1, seed=42,
)
print(f"Lag-1 fitness autocorrelation: {walk_autocorrelation:.4f}")
Output
Lag-1 fitness autocorrelation: 0.6915
6.3 Walsh–Hadamard decomposition¶
A metal substitution may affect the gap differently with different halides. walsh_hadamard() fits individual-constituent effects and interactions between pairs. Compare first- and second-order training r2; delta_r2 shows the additional variation explained by pairwise terms.
Three-way effects remain outside the fit. Coefficients describe the original band-gap scale even though we minimise it. The model’s uniform-product variance spectrum is distinct from the incremental training fit and does not establish predictive accuracy.
wh = analysis.walsh_hadamard(landscape, max_order=2)
wh["order_summary"]
Output
| order | r2 | delta_r2 | rmse | n_terms | rank | alpha | model_variance_fraction | n_nonzero | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 0.000000 | 0.000000 | 1.136369 | 1 | 1 | NaN | 0.000000 | <NA> |
| 1 | 1 | 0.943964 | 0.943964 | 0.269000 | 21 | 21 | NaN | 0.958162 | <NA> |
| 2 | 2 | 0.985182 | 0.041218 | 0.138328 | 102 | 102 | NaN | 0.041838 | <NA> |
We can also inspect the largest fitted nonconstant coefficients. Their units follow the response, and their signs depend on the encoded contrasts:
wh["coefficients"].query("order > 0").sort_values(
"coefficient", key=abs, ascending=False,
).head(8)
Output
| order | positions | term | coefficient | |
|---|---|---|---|---|
| 3 | 1 | (3,) | F_3_I | -2.745577 |
| 1 | 1 | (3,) | F_3_Br | -2.213302 |
| 2 | 1 | (3,) | F_3_Cl | -1.610529 |
| 70 | 2 | (1, 2) | ethylammonium_1_hydroxylammonium-Ge_2_Pb | 1.246250 |
| 94 | 2 | (1, 3) | ethylammonium_1_tetramethylammonium-F_3_I | -1.176033 |
| 65 | 2 | (1, 2) | ethylammonium_1_hydrazinium-Ge_2_Pb | 1.071975 |
| 35 | 2 | (1, 2) | ethylammonium_1_ammonium-Ge_2_Pb | 1.007050 |
| 19 | 1 | (1,) | ethylammonium_1_tetramethylammonium | 0.974258 |
6.4 Fitness-distance correlation¶
Distance counts how many constituent identities differ from the selected best formula. It does not encode structural or electronic similarity.
We use Spearman fdc to compare ranks. A positive value means lower responses tend to occur nearer the optimum. A value near zero indicates little monotonic association. This trend does not guarantee an improving path from every configuration.
fitness_distance_r = analysis.fdc(landscape, method="spearman")
print(f"Fitness-distance correlation: {fitness_distance_r:.4f}")
Output
Fitness-distance correlation: 0.5994
This graph has one local optimum. The additive fit already explains about 94.4% of response variation, increasing to 98.5% with pairwise terms. Positive FDC (0.599) matches the minimisation objective.
7. Running several analyses together¶
We can collect autocorrelation and FDC with analysis.profile(), keeping the same walk settings. landscape.n_lo supplies the count directly; the W–H result keeps its separate coefficient and order-summary tables.
analysis.profile(
landscape,
metrics=["autocorrelation", "fdc"],
params={"autocorrelation": {"walk_length": 20, "walk_times": 200, "lag": 1}},
seed=42, progress=False,
)
Output
autocorrelation 0.691489
fdc 0.599377
dtype: float64
8. Analysis reference¶
For further exploration, analysis.list_metrics() lists the metrics available to profile(). The worked W–H result also exposes coefficients and fit_info; its order summary separates cumulative training fit from the fitted model’s variance spectrum.
Data sources¶
Kim C. et al. (2017). A hybrid organic-inorganic perovskite dataset. This tutorial uses the Olympus perovskites lookup, containing one HSE06 gap per formula, rather than all 1,346 relaxed structures in the original study.