Gompertz Linear Units: A Smarter Building Block for Artificial Intelligence
Principal Investigator:
Prof. Dr. Frank Hutter
Affiliation:
Universität Freiburg, Department of Computer Science, Freiburg, Germany
Local Project ID:
pn68xi
HPC Platform used:
SuperMUC-NG PH1-CPU and SuperMUC-NG PH2-GPU
Date published:
Deep learning models power everything from image recognition to language translation, but their performance depends heavily on small internal components called activation functions. Researchers at the University of Freiburg have developed a new activation function called GoLU, the Gompertz Linear Unit, that consistently outperforms existing alternatives across a wide range of tasks. Using high performance computing clusters, the team conducted over 112,000 GPU hours of experiments, demonstrating that this deceptively simple mathematical function can make neural networks more accurate and robust.
Deep neural networks depend on activation functions, mathematical rules that determine how signals propagate through the network, yet the choice of activation function remains an open research question. The widely used ReLU function and its more recent successors like GELU and Swish each come with trade-offs, motivating the search for a better alternative that could improve learning across a broad range of applications without adding computational cost.
The researchers drew inspiration from mathematical ideas that trace back to Benjamin Gompertz's work on human mortality in the 19th century. In particular, they focused on the Gumbel distribution, whose asymmetric shape (Figure 1) differs from the more balanced curves that underlie many activation functions used in modern AI systems. Building on this idea, the team developed GoLU, a new activation function that uses this asymmetry to regulate the flow of information through a neural network. Figure 2 compares GoLU with widely used activation functions such as ReLU, GELU, and Swish. By selectively strengthening useful signals while suppressing less relevant ones, GoLU helps neural networks learn more effectively and reliably across a variety of tasks.
To convincingly demonstrate that a new activation function works better than established ones, it is not enough to test it on a single task or dataset. The team needed to run hundreds of experiments across very different domains, from training image classifiers on the ImageNet dataset (containing over a million photographs) to training language models on billions of words of text. Each individual experiment required training a full neural network from scratch, on powerful graphics processing units (GPUs). See Table 1 for compute time on HPC clusters.
The results were remarkably consistent. GoLU outperformed all tested alternatives across every task shown in Table 2. Notably, GELU, currently the default activation in many widely used models, was typically the second-best performer, but GoLU achieved measurable improvements over it in nearly every setting.
Beyond raw accuracy, the team discovered that GoLU produces smoother loss landscapes, meaning the optimization process is less sensitive to small changes in network parameters. This suggests that models trained with GoLU are not only more accurate but also more robust and easier to train. Furthermore, the researchers developed a custom CUDA kernel for GoLU, ensuring that its computational cost is virtually identical to that of existing activation functions.
Because activation functions are a universal component of virtually every neural network, the potential impact of GoLU is broad. Any practitioner training deep learning models, whether for medical imaging, autonomous driving or natural language processing, could benefit from simply replacing their current activation function with GoLU. The team has made their code and the optimized CUDA implementation publicly available to facilitate adoption by the research community and industry alike.
[1] Indrashis Das, Mahmoud Safari, Steven Adriaensen, Frank Hutter. "Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics." https://arxiv.org/abs/2502.03654. Advances in Neural Information Processing Systems. 2026 Apr 23;38:63699-730.
[2] https://github.com/automl/GoLU

Figure 1: The Gumbel distribution (red) underlying GoLU is visibly skewed to the right.

Figure 2: This right-skewed asymmetry causes GoLU (red) to stay closest to the x-axis, compressing internal signals and reducing noise.

Table 1: Summary of representative experiments conducted on HPC clusters.Total compute across all tasks and configurations: approximately 112,000 GPU hours.

Table 2: Summary of representative experiments. GoLU outperforms the best baseline activation function. ↓ = lower is better.