Neural architecture search¶
NAS-Bench-201 · CIFAR-10
In this tutorial, we’ll use GraphFLA to explore how neural-network architecture choices affect validation accuracy. To analyse these choices without training new networks, we’ll use recorded results from NAS-Bench-201 to build and examine an architecture landscape.
Download notebook Notebook + data
Run the notebook in Google Colab, or download it to run locally. It downloads its data when no data/ folder is beside it.
1. Setting up¶
On Google Colab, the cell below installs GraphFLA. In any environment, it also downloads the dataset into a data/ folder next to the notebook if the files are not already there.
We then import GraphFLA for landscape construction and analysis, and pandas for working with the data.
import sys
from pathlib import Path
from urllib.request import urlretrieve
if "google.colab" in sys.modules:
%pip install -q graphfla==0.4.0
DATA_URL = "https://raw.githubusercontent.com/COLA-Laboratory/GraphFLA/v0.4.0/tutorials/datasets/data/"
Path("data").mkdir(exist_ok=True)
for name in ["nasbench201.csv"]:
if not Path("data", name).exists():
urlretrieve(DATA_URL + name, Path("data", name))
from graphfla import analysis
from graphfla.landscape import Landscape
import pandas as pd
2. Loading the dataset¶
Neural architecture search compares network designs to find ones that perform well. To study such a search, we’ll use NAS-Bench-201, a benchmark that records training results for every design in a fixed space. Its designs vary a small building block called a cell, with four nodes and six directed edges.
Each edge chooses no connection, a skip connection, a 1 × 1 or 3 × 3 convolution, or 3 × 3 average pooling. These choices give 5⁶ = 15,625 labelled architectures. We’ll maximise CIFAR-10 validation accuracy after 200 epochs, using the same training seed (888) for every design.
| Column | Meaning |
|---|---|
edge_i_to_j |
Operation on the directed cell edge from node i to node j. |
objective_validation_accuracy |
Top-1 validation accuracy (%), at epoch index 199. |
Let’s load the cifar10-valid / x-valid results. Seed 888 covers every architecture; distinct labelled assignments remain separate even if their computations are equivalent.
df = pd.read_csv("data/nasbench201.csv", keep_default_na=False)
df[["architecture_id", "architecture_string", "objective_validation_accuracy"]].head()
Output
| architecture_id | architecture_string | objective_validation_accuracy | |
|---|---|---|---|
| 0 | 0 | |avg_pool_3x3~0|+|nor_conv_1x1~0|skip_connect~... | 81.616 |
| 1 | 1 | |nor_conv_3x3~0|+|nor_conv_3x3~0|avg_pool_3x3~... | 90.716 |
| 2 | 2 | |avg_pool_3x3~0|+|nor_conv_3x3~0|nor_conv_3x3~... | 81.856 |
| 3 | 3 | |avg_pool_3x3~0|+|skip_connect~0|none~1|+|none... | 83.908 |
| 4 | 4 | |skip_connect~0|+|skip_connect~0|nor_conv_1x1~... | 85.148 |
3. Preparing the inputs¶
To construct a landscape, GraphFLA needs two aligned inputs:
X: one row per configuration and one column per variable.f: one measured or calculated outcome for each row ofX.
The six edge operations form X, in the order 0→1, 0→2, 1→2, 0→3, 1→3 and 2→3. Accuracy forms f. Training budget, dataset and seed are fixed background conditions rather than search variables.
edge_columns = [
"edge_0_to_1", "edge_0_to_2", "edge_1_to_2",
"edge_0_to_3", "edge_1_to_3", "edge_2_to_3",
]
X = df[edge_columns]
f = df["objective_validation_accuracy"]
X.head()
Output
| edge_0_to_1 | edge_0_to_2 | edge_1_to_2 | edge_0_to_3 | edge_1_to_3 | edge_2_to_3 | |
|---|---|---|---|---|---|---|
| 0 | avg_pool_3x3 | nor_conv_1x1 | skip_connect | nor_conv_1x1 | skip_connect | skip_connect |
| 1 | nor_conv_3x3 | nor_conv_3x3 | avg_pool_3x3 | skip_connect | nor_conv_3x3 | skip_connect |
| 2 | avg_pool_3x3 | nor_conv_3x3 | nor_conv_3x3 | avg_pool_3x3 | avg_pool_3x3 | avg_pool_3x3 |
| 3 | avg_pool_3x3 | skip_connect | none | none | none | skip_connect |
| 4 | skip_connect | skip_connect | nor_conv_1x1 | skip_connect | skip_connect | nor_conv_1x1 |
4. Constructing the landscape¶
We can use Landscape with six categorical variables. Neighbours replace the operation on one cell edge while leaving all other choices fixed. The none operation is a valid architectural choice.
We set maximize=True to favour higher validation accuracy. This analyses the recorded single-seed outcomes; it does not average away training variability.
landscape = Landscape(maximize=True)
landscape.build_from_data(
X, f, data_types={column: "categorical" for column in X},
neighborhood_strategy="active", verbose=False,
)
Output
Landscape(kind='default', maximize=True)
5. Inspecting the landscape¶
Let’s first look at what we’ve built. Printing the landscape shows its variables, configurations, improving edges and local optima:
print(landscape)
Output
Landscape(kind='default'): 6 variables, 15625 configurations, 186458 edges, 32 local optima
For a closer look at individual cell architectures, we can call get_data(). The outcome appears as fitness; out_degree counts directly improving moves, and is_lo indicates local-optimum membership.
landscape.get_data().head()
Output
| edge_0_to_1 | edge_0_to_2 | edge_1_to_2 | edge_0_to_3 | edge_1_to_3 | edge_2_to_3 | fitness | plateau_id | plateau_size | in_degree | out_degree | is_lo | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | avg_pool_3x3 | nor_conv_1x1 | skip_connect | nor_conv_1x1 | skip_connect | skip_connect | 81.616 | -1 | 1 | 7 | 17 | False |
| 1 | nor_conv_3x3 | nor_conv_3x3 | avg_pool_3x3 | skip_connect | nor_conv_3x3 | skip_connect | 90.716 | -1 | 1 | 15 | 9 | False |
| 2 | avg_pool_3x3 | nor_conv_3x3 | nor_conv_3x3 | avg_pool_3x3 | avg_pool_3x3 | avg_pool_3x3 | 81.856 | -1 | 1 | 8 | 16 | False |
| 3 | avg_pool_3x3 | skip_connect | none | none | none | skip_connect | 83.908 | -1 | 1 | 10 | 14 | False |
| 4 | skip_connect | skip_connect | nor_conv_1x1 | skip_connect | skip_connect | nor_conv_1x1 | 85.148 | -1 | 1 | 11 | 13 | False |
We can also summarise the validation accuracies with pandas:
landscape.get_data()["fitness"].describe()
Output
count 15625.000000
mean 83.629758
std 12.830918
min 9.712000
25% 84.288000
50% 87.244000
75% 88.940000
max 91.720000
Name: fitness, dtype: float64
To inspect the architecture with the highest recorded validation accuracy, we can use the global-optimum node index:
landscape[landscape.go_index]
Output
{'edge_0_to_1': 'nor_conv_3x3',
'edge_0_to_2': 'nor_conv_3x3',
'edge_1_to_2': 'nor_conv_3x3',
'edge_0_to_3': 'skip_connect',
'edge_1_to_3': 'nor_conv_3x3',
'edge_2_to_3': 'nor_conv_1x1',
'fitness': np.float64(91.71999999023436),
'plateau_id': -1,
'plateau_size': 1,
'in_degree': 24,
'out_degree': 0,
'is_lo': True}
6. Analysing the landscape¶
Next, we’ll count local optima, examine local fitness similarity and interaction orders, and assess the trend towards the best observed configuration.
6.1 Number of local optima¶
We can start with landscape.n_lo. Each connected neutral optimum plateau counts once. A local optimum has no improving move out of it, including through its neutral plateau. Multiple optima mean successive improvements can end at different configurations.
print(f"Number of local optima: {landscape.n_lo}")
Output
Number of local optima: 32
6.2 Fitness autocorrelation¶
How similar are responses at neighbouring steps? autocorrelation() measures persistence along sampled walks. Higher values indicate more persistence under the same walk settings.
We use 200 walks of up to 20 visited configurations, lag one and a fixed seed. Walks traverse stored edges in either direction; separately stored neutral pairs are excluded.
walk_autocorrelation = analysis.autocorrelation(
landscape, walk_length=20, walk_times=200, lag=1, seed=42,
)
print(f"Lag-1 fitness autocorrelation: {walk_autocorrelation:.4f}")
Output
Lag-1 fitness autocorrelation: 0.4594
6.3 Walsh–Hadamard decomposition¶
An operation’s contribution may depend on choices on other edges. walsh_hadamard() fits individual-edge effects and pairwise interactions. Comparing the first- and second-order r2 values shows how much training variation the pairwise terms add.
Interactions involving three or more edges remain outside this fit. model_variance_fraction describes the fitted model under uniform operation assignments, whereas delta_r2 measures improvement between nested fits. Neither quantity is an out-of-sample prediction score.
wh = analysis.walsh_hadamard(landscape, max_order=2)
wh["order_summary"]
Output
| order | r2 | delta_r2 | rmse | n_terms | rank | alpha | model_variance_fraction | n_nonzero | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 0.000000 | 0.000000 | 12.830508 | 1 | 1 | NaN | 0.000000 | <NA> |
| 1 | 1 | 0.291169 | 0.291169 | 10.802272 | 25 | 25 | NaN | 0.425114 | <NA> |
| 2 | 2 | 0.684921 | 0.393752 | 7.202003 | 265 | 265 | NaN | 0.574886 | <NA> |
We can also inspect the largest fitted nonconstant coefficients. Their units follow the response, and their signs depend on the encoded contrasts:
wh["coefficients"].query("order > 0").sort_values(
"coefficient", key=abs, ascending=False,
).head(8)
Output
| order | positions | term | coefficient | |
|---|---|---|---|---|
| 182 | 2 | (4, 6) | nor%5Fconv%5F1x1_4_none-skip%5Fconnect_6_none | -25.118829 |
| 30 | 2 | (1, 4) | avg%5Fpool%5F3x3_1_none-nor%5Fconv%5F1x1_4_none | -20.919494 |
| 122 | 2 | (2, 4) | nor%5Fconv%5F1x1_2_none-nor%5Fconv%5F1x1_4_none | -15.777715 |
| 254 | 2 | (5, 6) | skip%5Fconnect_5_none-skip%5Fconnect_6_none | -14.933331 |
| 178 | 2 | (4, 5) | nor%5Fconv%5F1x1_4_none-skip%5Fconnect_5_none | -14.566394 |
| 10 | 1 | (4,) | nor%5Fconv%5F1x1_4_none | -11.385916 |
| 42 | 2 | (1, 6) | avg%5Fpool%5F3x3_1_none-skip%5Fconnect_6_none | -11.240288 |
| 26 | 2 | (1, 2) | avg%5Fpool%5F3x3_1_none-nor%5Fconv%5F1x1_2_none | -10.949018 |
6.4 Fitness-distance correlation¶
Distance counts cell edges whose operations differ from the selected best architecture. It does not measure differences in parameter count, latency or computational cost.
We use Spearman fdc to compare ranks. A negative value means higher responses tend to occur nearer the optimum. A value near zero indicates little monotonic association. This trend does not guarantee an improving path from every configuration.
fitness_distance_r = analysis.fdc(landscape, method="spearman")
print(f"Fitness-distance correlation: {fitness_distance_r:.4f}")
Output
Fitness-distance correlation: -0.4132
Pairwise terms raise training R² from 0.291 to 0.685. The remaining variation is unresolved by this truncated model. The 32 local optima also show why successive single-edge improvements can stop at different architectures.
7. Running several analyses together¶
We can collect autocorrelation and FDC with analysis.profile(), keeping the same walk settings. landscape.n_lo supplies the count directly; the W–H result keeps its separate coefficient and order-summary tables.
analysis.profile(
landscape,
metrics=["autocorrelation", "fdc"],
params={"autocorrelation": {"walk_length": 20, "walk_times": 200, "lag": 1}},
seed=42, progress=False,
)
Output
autocorrelation 0.459357
fdc -0.413214
dtype: float64
8. Analysis reference¶
For further exploration, analysis.list_metrics() lists the metrics available to profile(). The worked W–H result also exposes coefficients and fit_info; its order summary separates cumulative training fit from the fitted model’s variance spectrum.
Data sources¶
Dong X. and Yang Y. (2020). NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture Search. The compact CSV was extracted from the authors’ official v1.0 benchmark archive using one fixed, complete training seed. See the author repository.