Alloy composition optimization¶
Tungsten–rhenium–osmium alloys
In this tutorial, we’ll use GraphFLA to explore how an alloy’s composition affects its hardness at high temperatures. To connect composition choices with measured performance, we’ll build and analyse a landscape of tungsten–rhenium–osmium alloys.
Download notebook Notebook + data
Run the notebook in Google Colab, or download it to run locally. It downloads its data when no data/ folder is beside it.
1. Setting up¶
On Google Colab, the cell below installs GraphFLA. In any environment, it also downloads the dataset into a data/ folder next to the notebook if the files are not already there.
We then import GraphFLA for landscape construction and analysis, and pandas for working with the data.
import sys
from pathlib import Path
from urllib.request import urlretrieve
if "google.colab" in sys.modules:
%pip install -q graphfla==0.4.0
DATA_URL = "https://raw.githubusercontent.com/COLA-Laboratory/GraphFLA/v0.4.0/tutorials/datasets/data/"
Path("data").mkdir(exist_ok=True)
for name in ["wreos.csv"]:
if not Path("data", name).exists():
urlretrieve(DATA_URL + name, Path("data", name))
from graphfla import analysis
from graphfla.landscape import OrdinalLandscape
import pandas as pd
2. Loading the dataset¶
Developing alloys for high-temperature applications involves choosing proportions that retain useful mechanical properties when heated. Wang et al. (2026) investigated this problem by producing and testing many mixtures of tungsten (W), rhenium (Re) and osmium (Os).
To explore their composition space, we’ll use 496 measured alloys and maximise Vickers hardness at 1000 °C. W and Re vary in three-atomic-percent steps, with Os filling the balance. Each row gives a composition and its measured properties:
| Column | Meaning |
|---|---|
W, Re, Os |
Nominal atomic percentages, summing to 100. |
H1000_HV |
Vickers hardness at 1000 °C. |
Let’s load the alloy measurements. The three pure-element controls are outside this grid; we’ll focus on hardness, leaving the study’s other mechanical properties aside.
df = pd.read_csv("data/wreos.csv")
df.head()
Output
| composition_id | W | Re | Os | H25_HV | H1000_HV | R1000 | pileup_um | TOPSIS_Ci | phase_label | |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | W3Re3Os94 | 3 | 3 | 94 | 741 | 411 | 0.55 | 0.1 | 0.22 | HCP |
| 1 | W3Re6Os91 | 3 | 6 | 91 | 793 | 423 | 0.53 | 0.4 | 0.21 | HCP |
| 2 | W3Re9Os88 | 3 | 9 | 88 | 794 | 442 | 0.56 | 0.0 | 0.23 | HCP |
| 3 | W3Re12Os85 | 3 | 12 | 85 | 794 | 460 | 0.58 | 0.0 | 0.26 | HCP |
| 4 | W3Re15Os82 | 3 | 15 | 82 | 768 | 454 | 0.59 | 0.0 | 0.27 | HCP |
3. Preparing the inputs¶
To construct a landscape, GraphFLA needs two aligned inputs:
X: one row per configuration and one column per variable.f: one measured or calculated outcome for each row ofX.
W and Re each take values 3, 6, …, 93, subject to W + Re <= 96. This triangular grid has 496 feasible combinations, rather than 31 × 31 independent choices. We leave Os out of X because Os = 100 - W - Re.
X = df[["W", "Re"]]
f = df["H1000_HV"]
X.head()
Output
| W | Re | |
|---|---|---|
| 0 | 3 | 3 |
| 1 | 3 | 6 |
| 2 | 3 | 9 |
| 3 | 3 | 12 |
| 4 | 3 | 15 |
4. Constructing the landscape¶
We can use OrdinalLandscape because the sampled composition levels are ordered. Neighbours change W or Re by one sampled level, holding the other coordinate fixed. Os adjusts to preserve the total.
Each move exchanges 3 atomic percent with Os; a direct W–Re exchange requires two moves. Missing combinations outside the triangle are infeasible. We set maximize=True to favour higher hardness.
landscape = OrdinalLandscape(maximize=True)
landscape.build_from_data(X, f, verbose=False)
Output
OrdinalLandscape(maximize=True)
5. Inspecting the landscape¶
Let’s first look at what we’ve built. Printing the landscape shows its variables, configurations, improving edges and local optima:
print(landscape)
Output
OrdinalLandscape(kind='ordinal'): 2 variables, 496 configurations, 911 edges, 36 local optima
For a closer look at individual alloy compositions, we can call get_data(). The outcome appears as fitness; out_degree counts directly improving moves, and is_lo indicates local-optimum membership.
landscape.get_data().head()
Output
| W | Re | fitness | plateau_id | plateau_size | in_degree | out_degree | is_lo | |
|---|---|---|---|---|---|---|---|---|
| 0 | 3 | 3 | 411.0 | -1 | 1 | 0 | 2 | False |
| 1 | 3 | 6 | 423.0 | -1 | 1 | 1 | 2 | False |
| 2 | 3 | 9 | 442.0 | -1 | 1 | 1 | 2 | False |
| 3 | 3 | 12 | 460.0 | -1 | 1 | 2 | 1 | False |
| 4 | 3 | 15 | 454.0 | -1 | 1 | 1 | 2 | False |
We can also summarise the hardness measurements with pandas:
landscape.get_data()["fitness"].describe()
Output
count 496.000000
mean 609.294355
std 296.659404
min 201.000000
25% 396.000000
50% 453.500000
75% 826.750000
max 1265.000000
Name: fitness, dtype: float64
To inspect the composition with the highest measured hardness, we can use the global-optimum node index:
landscape[landscape.go_index]
Output
{'W': np.int64(63),
'Re': np.int64(21),
'fitness': np.float64(1265.0),
'plateau_id': -1,
'plateau_size': 1,
'in_degree': 4,
'out_degree': 0,
'is_lo': True}
6. Analysing the landscape¶
Next, we’ll count local optima, examine local fitness similarity and interaction orders, and assess the trend towards the best observed configuration.
6.1 Number of local optima¶
We can start with landscape.n_lo. Each connected neutral optimum plateau counts once. A local optimum has no improving move out of it, including through its neutral plateau. Multiple optima mean successive improvements can end at different configurations.
print(f"Number of local optima: {landscape.n_lo}")
Output
Number of local optima: 36
6.2 Fitness autocorrelation¶
How similar are responses at neighbouring steps? autocorrelation() measures persistence along sampled walks. Higher values indicate more persistence under the same walk settings.
We use 200 walks of up to 20 visited configurations, lag one and a fixed seed. Walks traverse stored edges in either direction; separately stored neutral pairs are excluded.
walk_autocorrelation = analysis.autocorrelation(
landscape, walk_length=20, walk_times=200, lag=1, seed=42,
)
print(f"Lag-1 fitness autocorrelation: {walk_autocorrelation:.4f}")
Output
Lag-1 fitness autocorrelation: 0.9366
6.3 Walsh–Hadamard decomposition¶
walsh_hadamard() treats each sampled composition level as a discrete state. We fit through order two using explicit Lasso regularisation: the triangular grid cannot identify every coefficient of the full 31 × 31 model by ordinary least squares.
Compare cumulative training r2 across orders. Cross-validation selects each penalty, but these displayed scores are not held-out performance. The model’s full-product variance spectrum extends to unobserved, infeasible combinations, so we do not interpret it as measured variance on the alloy triangle.
wh = analysis.walsh_hadamard(
landscape, max_order=2, method="lasso", alpha="cv",
cv=5, random_state=42, n_jobs=1,
)
wh["order_summary"][["order", "r2", "delta_r2", "rmse", "alpha", "n_nonzero"]]
Output
| order | r2 | delta_r2 | rmse | alpha | n_nonzero | |
|---|---|---|---|---|---|---|
| 0 | 0 | 0.000000 | 0.000000 | 296.360201 | NaN | 1 |
| 1 | 1 | 0.911758 | 0.911758 | 88.035591 | 0.014064 | 55 |
| 2 | 2 | 0.912804 | 0.001046 | 87.512294 | 0.348380 | 29 |
We can also inspect the largest fitted nonconstant coefficients. Their units follow the response, and their signs depend on the encoded contrasts:
wh["coefficients"].query("order > 0").sort_values(
"coefficient", key=abs, ascending=False,
).head(8)
Output
| order | positions | term | coefficient | |
|---|---|---|---|---|
| 20 | 1 | (1,) | 3_1_66 | 703.222769 |
| 18 | 1 | (1,) | 3_1_60 | 691.978524 |
| 19 | 1 | (1,) | 3_1_63 | 686.128151 |
| 21 | 1 | (1,) | 3_1_69 | 664.597298 |
| 16 | 1 | (1,) | 3_1_57 | 657.865363 |
| 15 | 1 | (1,) | 3_1_54 | 595.553662 |
| 22 | 1 | (1,) | 3_1_72 | 595.173613 |
| 14 | 1 | (1,) | 3_1_51 | 532.244389 |
6.4 Fitness-distance correlation¶
Distance sums the sampled W and Re level differences. It follows the chosen two-coordinate geometry, with Os as the balance element.
We use Spearman fdc to compare ranks. A negative value means higher responses tend to occur nearer the optimum. A value near zero indicates little monotonic association. This trend does not guarantee an improving path from every configuration.
fitness_distance_r = analysis.fdc(landscape, method="spearman")
print(f"Fitness-distance correlation: {fitness_distance_r:.4f}")
Output
Fitness-distance correlation: -0.8687
The 36 local optima coexist with strong autocorrelation (0.937) and FDC of −0.869. Local peaks can therefore remain within a landscape that has a clear overall trend towards high hardness.
7. Running several analyses together¶
We can collect autocorrelation and FDC with analysis.profile(), keeping the same walk settings. landscape.n_lo supplies the count directly; the W–H result keeps its separate coefficient and order-summary tables.
analysis.profile(
landscape,
metrics=["autocorrelation", "fdc"],
params={"autocorrelation": {"walk_length": 20, "walk_times": 200, "lag": 1}},
seed=42, progress=False,
)
Output
autocorrelation 0.936634
fdc -0.868738
dtype: float64
8. Analysis reference¶
For further exploration, analysis.list_metrics() lists the metrics available to profile(). The worked W–H result also exposes coefficients and fit_info; its order summary separates cumulative training fit from the fitted model’s variance spectrum.
Data sources¶
Wang et al. (2026). High-throughput discovery of ultrahigh-temperature multi-principal element alloys by combinatorial additive manufacturing. The prepared table agrees with the article’s Source Data.