A Next Generation Benchmark for Automated Deep Learning
Principal Investigator:
Prof. Dr. Frank Hutter
Affiliation:
Universität Freiburg, Department of Computer Science, Freiburg, Germany
Local Project ID:
pn68xi
HPC Platform used:
SuperMUC-NG PH1-CPU and SuperMUC-NG PH2-GPU
Date published:
Since the release of ChatGPT in December 2022, artificial intelligence (AI) has garnered unprecedented media attention. Mainstream media outlets, such as Tagesschau, now regularly report AI-related news, and public discourse on the subject has permeated schools and universities. This widespread attention is the result of significant developments in Deep Learning (DL), which has proven incredibly successful at solving previously unsolvable tasks and outperforming known solutions to established problems. Deep Neural Networks (DNNs) are very powerful, but designing their architectures and training pipelines for novel tasks has traditionally required extensive trial-and-error by human experts. Thus, DNNs have been hard to adopt in new fields.
To democratise DL and enable inexperienced users to design good DNNs for novel tasks without the need for extensive research into domain-specific DNN model designs, NAS (Neural Architecture Search) treats the design of these network architectures itself as an optimization problem that can be solved algorithmically. NAS has recently shown very promising results in designing novel DNNs that outperform hand-crafted ones. Parallel to NAS, research in automating the training pipelines for DNNs focuses on algorithmic approaches towards choosing the correct values of a number of parameters, called hyperparameters, that alter the behaviour of the model training pipeline. This field of research is called Hyperparameter Optimization (HPO).
Research on both NAS and HPO is extremely compute intensive, since the evaluation of a single NAS or HPO algorithm may involve training hundreds or thousands of DNNs, cumulatively consuming several weeks or months of compute on large compute clusters. This presents a high barrier to entry for researchers who do not have access to such amounts of compute, as well as raises concerns about the carbon footprint of such research. In response, PI Frank Hutter has previously pioneered the development of benchmarks as a one-time investment of resources in order to amortise the costs of all future NAS research, starting a trend amongst NAS researchers of developing a diverse range of benchmarks.
As the impact of architectures and hyperparameters is intrinsically linked, a joint approach towards optimising both architectures and hyperparameters can be more powerful than NAS or HPO alone. Nevertheless, just as NAS and HPO research was previously bottlenecked by an absence of appropriate benchmarks, so too is research into this problem currently blocked by the absence of suitable benchmarks. Therefore, we proposed JAHS-Bench-201 to enable research into what we refer to as Joint Architecture and Hyperparameter Search (JAHS). JAHS-Bench-201 is likely to become the cornerstone for future research in as-of-yet underexplored approaches of JAHS, and it has already gained significant recognition for its contributions as it was accepted as a featured paper (i.e. top-7.5% of accepted papers and top-2% of submitted papers) in the prestigious 36th Conference on Neural Information Processing Systems’ (NeurIPS’22) Datasets and Benchmarks Track [1].
To create JAHS-Bench-201 [2], we designed a 14-dimensional search space and used the SuperMUC-NG cluster to collect the most extensive dataset of neural network performance data available in the public domain. We used this performance dataset to train surrogate models, so researchers can now use these models to query the JAHS-Bench-201 search space for an approximate evaluation without actually training a complete DNN. This technique reduces the associated compute time from GPU-days to CPU-seconds (see Table 1). A single research paper may need to perform many hundreds of thousands of such evaluations, therefore using our benchmark will both help democratise research on JAHS and save substantial compute time & carbon emissions.
For our search space, we chose to combine one of the most popular search spaces in NAS literature - the search space of the benchmark NAS-Bench-201 - with some common choices for hyperparameters in the HPO literature. We complement the main search space with fidelity parameters that represent low-cost approximations such as reduced network size or reduced image resolutions, as NAS and HPO algorithms use these to gain speedups. As we include cheap low-fidelity evaluations, in contrast to the predominant use of GPUs in other areas of deep learning, our experiments are actually much more cost-efficient to perform on a large cluster of CPUs; thus, we carried them out on SuperMUC-NG.
Our performance dataset is composed of approximately 161 million data points, each consisting of 20 performance metrics, across three image classification tasks. As we record the metrics at each epoch of model training, we yield an additional fidelity parameter in the form of the number of epochs of training. Therefore, JAHS-Bench-201 supports not just 4 distinct fidelity parameters, but also supports research into the joint impact of varying multiple fidelity parameters simultaneously, which is unprecedented in extant related literature.
We were able to utilise the massive number of parallel compute nodes available on the SuperMUC-NG to run simulations for up to 100k Deep Learning trainings in parallel. The greatest limitation we faced was the availability of RAM on a per-node basis, which bottlenecked the number of parallel computations that we can perform. Nonetheless, the SuperMUC-NG provided us with parallel compute resources well beyond what we see in related work (see Table 2).
[1] A. Bansal, D. Stoll, M. Janowski, A. Zela, and F. Hutter. JAHS-Bench-201: A Foundation For Research On Joint Architecture And Hyperparameter Search. In Proceedings of the Thirty-sixth Conference on Neural Information Processing Systems. 2022.
[2] https://github.com/automl/jahs_bench_201

Table 1: Comparison of the compute requirements of surrogate-based experiments (values reported in CPU-seconds) vs training-based experiments (values reported in GPU-days).

Table 2: Comparison of JAHS-Bench-201 to NAS-HPO-Bench-II (October, 2021), which shares a number of common properties with our work but is much more limited in its scope.