Skip to content

Neural architecture search

NAS-Bench-201 · CIFAR-10

In this tutorial, we’ll use GraphFLA to explore how neural-network architecture choices affect validation accuracy. To analyse these choices without training new networks, we’ll use recorded results from NAS-Bench-201 to build and examine an architecture landscape.

Open In Colab

Download notebook Notebook + data

Run the notebook in Google Colab, or download it to run locally. It downloads its data when no data/ folder is beside it.

1. Setting up

On Google Colab, the cell below installs GraphFLA. In any environment, it also downloads the dataset into a data/ folder next to the notebook if the files are not already there.

We then import GraphFLA for landscape construction and analysis, and pandas for working with the data.

import sys
from pathlib import Path
from urllib.request import urlretrieve

if "google.colab" in sys.modules:
    %pip install -q graphfla==0.4.0

DATA_URL = "https://raw.githubusercontent.com/COLA-Laboratory/GraphFLA/v0.4.0/tutorials/datasets/data/"
Path("data").mkdir(exist_ok=True)
for name in ["nasbench201.csv"]:
    if not Path("data", name).exists():
        urlretrieve(DATA_URL + name, Path("data", name))
from graphfla import analysis
from graphfla.landscape import Landscape
import pandas as pd

2. Loading the dataset

Neural architecture search compares network designs to find ones that perform well. To study such a search, we’ll use NAS-Bench-201, a benchmark that records training results for every design in a fixed space. Its designs vary a small building block called a cell, with four nodes and six directed edges.

Each edge chooses no connection, a skip connection, a 1 × 1 or 3 × 3 convolution, or 3 × 3 average pooling. These choices give 5⁶ = 15,625 labelled architectures. We’ll maximise CIFAR-10 validation accuracy after 200 epochs, using the same training seed (888) for every design.

Column Meaning
edge_i_to_j Operation on the directed cell edge from node i to node j.
objective_validation_accuracy Top-1 validation accuracy (%), at epoch index 199.

Let’s load the cifar10-valid / x-valid results. Seed 888 covers every architecture; distinct labelled assignments remain separate even if their computations are equivalent.

df = pd.read_csv("data/nasbench201.csv", keep_default_na=False)
df[["architecture_id", "architecture_string", "objective_validation_accuracy"]].head()

Output

architecture_id architecture_string objective_validation_accuracy
0 0 |avg_pool_3x3~0|+|nor_conv_1x1~0|skip_connect~... 81.616
1 1 |nor_conv_3x3~0|+|nor_conv_3x3~0|avg_pool_3x3~... 90.716
2 2 |avg_pool_3x3~0|+|nor_conv_3x3~0|nor_conv_3x3~... 81.856
3 3 |avg_pool_3x3~0|+|skip_connect~0|none~1|+|none... 83.908
4 4 |skip_connect~0|+|skip_connect~0|nor_conv_1x1~... 85.148

3. Preparing the inputs

To construct a landscape, GraphFLA needs two aligned inputs:

  • X: one row per configuration and one column per variable.
  • f: one measured or calculated outcome for each row of X.

The six edge operations form X, in the order 0→1, 0→2, 1→2, 0→3, 1→3 and 2→3. Accuracy forms f. Training budget, dataset and seed are fixed background conditions rather than search variables.

edge_columns = [
    "edge_0_to_1", "edge_0_to_2", "edge_1_to_2",
    "edge_0_to_3", "edge_1_to_3", "edge_2_to_3",
]
X = df[edge_columns]
f = df["objective_validation_accuracy"]
X.head()

Output

edge_0_to_1 edge_0_to_2 edge_1_to_2 edge_0_to_3 edge_1_to_3 edge_2_to_3
0 avg_pool_3x3 nor_conv_1x1 skip_connect nor_conv_1x1 skip_connect skip_connect
1 nor_conv_3x3 nor_conv_3x3 avg_pool_3x3 skip_connect nor_conv_3x3 skip_connect
2 avg_pool_3x3 nor_conv_3x3 nor_conv_3x3 avg_pool_3x3 avg_pool_3x3 avg_pool_3x3
3 avg_pool_3x3 skip_connect none none none skip_connect
4 skip_connect skip_connect nor_conv_1x1 skip_connect skip_connect nor_conv_1x1

4. Constructing the landscape

We can use Landscape with six categorical variables. Neighbours replace the operation on one cell edge while leaving all other choices fixed. The none operation is a valid architectural choice.

We set maximize=True to favour higher validation accuracy. This analyses the recorded single-seed outcomes; it does not average away training variability.

landscape = Landscape(maximize=True)
landscape.build_from_data(
    X, f, data_types={column: "categorical" for column in X},
    neighborhood_strategy="active", verbose=False,
)

Output

Landscape(kind='default', maximize=True)

5. Inspecting the landscape

Let’s first look at what we’ve built. Printing the landscape shows its variables, configurations, improving edges and local optima:

print(landscape)

Output

Landscape(kind='default'): 6 variables, 15625 configurations, 186458 edges, 32 local optima

For a closer look at individual cell architectures, we can call get_data(). The outcome appears as fitness; out_degree counts directly improving moves, and is_lo indicates local-optimum membership.

landscape.get_data().head()

Output

edge_0_to_1 edge_0_to_2 edge_1_to_2 edge_0_to_3 edge_1_to_3 edge_2_to_3 fitness plateau_id plateau_size in_degree out_degree is_lo
0 avg_pool_3x3 nor_conv_1x1 skip_connect nor_conv_1x1 skip_connect skip_connect 81.616 -1 1 7 17 False
1 nor_conv_3x3 nor_conv_3x3 avg_pool_3x3 skip_connect nor_conv_3x3 skip_connect 90.716 -1 1 15 9 False
2 avg_pool_3x3 nor_conv_3x3 nor_conv_3x3 avg_pool_3x3 avg_pool_3x3 avg_pool_3x3 81.856 -1 1 8 16 False
3 avg_pool_3x3 skip_connect none none none skip_connect 83.908 -1 1 10 14 False
4 skip_connect skip_connect nor_conv_1x1 skip_connect skip_connect nor_conv_1x1 85.148 -1 1 11 13 False

We can also summarise the validation accuracies with pandas:

landscape.get_data()["fitness"].describe()

Output

count    15625.000000
mean        83.629758
std         12.830918
min          9.712000
25%         84.288000
50%         87.244000
75%         88.940000
max         91.720000
Name: fitness, dtype: float64

To inspect the architecture with the highest recorded validation accuracy, we can use the global-optimum node index:

landscape[landscape.go_index]

Output

{'edge_0_to_1': 'nor_conv_3x3',
 'edge_0_to_2': 'nor_conv_3x3',
 'edge_1_to_2': 'nor_conv_3x3',
 'edge_0_to_3': 'skip_connect',
 'edge_1_to_3': 'nor_conv_3x3',
 'edge_2_to_3': 'nor_conv_1x1',
 'fitness': np.float64(91.71999999023436),
 'plateau_id': -1,
 'plateau_size': 1,
 'in_degree': 24,
 'out_degree': 0,
 'is_lo': True}

6. Analysing the landscape

Next, we’ll count local optima, examine local fitness similarity and interaction orders, and assess the trend towards the best observed configuration.

6.1 Number of local optima

We can start with landscape.n_lo. Each connected neutral optimum plateau counts once. A local optimum has no improving move out of it, including through its neutral plateau. Multiple optima mean successive improvements can end at different configurations.

print(f"Number of local optima: {landscape.n_lo}")

Output

Number of local optima: 32

6.2 Fitness autocorrelation

How similar are responses at neighbouring steps? autocorrelation() measures persistence along sampled walks. Higher values indicate more persistence under the same walk settings.

We use 200 walks of up to 20 visited configurations, lag one and a fixed seed. Walks traverse stored edges in either direction; separately stored neutral pairs are excluded.

walk_autocorrelation = analysis.autocorrelation(
    landscape, walk_length=20, walk_times=200, lag=1, seed=42,
)
print(f"Lag-1 fitness autocorrelation: {walk_autocorrelation:.4f}")

Output

Lag-1 fitness autocorrelation: 0.4594

6.3 Walsh–Hadamard decomposition

An operation’s contribution may depend on choices on other edges. walsh_hadamard() fits individual-edge effects and pairwise interactions. Comparing the first- and second-order r2 values shows how much training variation the pairwise terms add.

Interactions involving three or more edges remain outside this fit. model_variance_fraction describes the fitted model under uniform operation assignments, whereas delta_r2 measures improvement between nested fits. Neither quantity is an out-of-sample prediction score.

wh = analysis.walsh_hadamard(landscape, max_order=2)
wh["order_summary"]

Output

order r2 delta_r2 rmse n_terms rank alpha model_variance_fraction n_nonzero
0 0 0.000000 0.000000 12.830508 1 1 NaN 0.000000 <NA>
1 1 0.291169 0.291169 10.802272 25 25 NaN 0.425114 <NA>
2 2 0.684921 0.393752 7.202003 265 265 NaN 0.574886 <NA>

We can also inspect the largest fitted nonconstant coefficients. Their units follow the response, and their signs depend on the encoded contrasts:

wh["coefficients"].query("order > 0").sort_values(
    "coefficient", key=abs, ascending=False,
).head(8)

Output

order positions term coefficient
182 2 (4, 6) nor%5Fconv%5F1x1_4_none-skip%5Fconnect_6_none -25.118829
30 2 (1, 4) avg%5Fpool%5F3x3_1_none-nor%5Fconv%5F1x1_4_none -20.919494
122 2 (2, 4) nor%5Fconv%5F1x1_2_none-nor%5Fconv%5F1x1_4_none -15.777715
254 2 (5, 6) skip%5Fconnect_5_none-skip%5Fconnect_6_none -14.933331
178 2 (4, 5) nor%5Fconv%5F1x1_4_none-skip%5Fconnect_5_none -14.566394
10 1 (4,) nor%5Fconv%5F1x1_4_none -11.385916
42 2 (1, 6) avg%5Fpool%5F3x3_1_none-skip%5Fconnect_6_none -11.240288
26 2 (1, 2) avg%5Fpool%5F3x3_1_none-nor%5Fconv%5F1x1_2_none -10.949018

6.4 Fitness-distance correlation

Distance counts cell edges whose operations differ from the selected best architecture. It does not measure differences in parameter count, latency or computational cost.

We use Spearman fdc to compare ranks. A negative value means higher responses tend to occur nearer the optimum. A value near zero indicates little monotonic association. This trend does not guarantee an improving path from every configuration.

fitness_distance_r = analysis.fdc(landscape, method="spearman")
print(f"Fitness-distance correlation: {fitness_distance_r:.4f}")

Output

Fitness-distance correlation: -0.4132

Pairwise terms raise training R² from 0.291 to 0.685. The remaining variation is unresolved by this truncated model. The 32 local optima also show why successive single-edge improvements can stop at different architectures.

7. Running several analyses together

We can collect autocorrelation and FDC with analysis.profile(), keeping the same walk settings. landscape.n_lo supplies the count directly; the W–H result keeps its separate coefficient and order-summary tables.

analysis.profile(
    landscape,
    metrics=["autocorrelation", "fdc"],
    params={"autocorrelation": {"walk_length": 20, "walk_times": 200, "lag": 1}},
    seed=42, progress=False,
)

Output

autocorrelation    0.459357
fdc               -0.413214
dtype: float64

8. Analysis reference

For further exploration, analysis.list_metrics() lists the metrics available to profile(). The worked W–H result also exposes coefficients and fit_info; its order summary separates cumulative training fit from the fitted model’s variance spectrum.

Data sources

Dong X. and Yang Y. (2020). NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture Search. The compact CSV was extracted from the authors’ official v1.0 benchmark archive using one fixed, complete training seed. See the author repository.