Exascale computing is not just another milestone in the history of supercomputers – it is a fundamental shift in how we approach scientific research, medicine, and engineering. Systems like JUPITER, Frontier, or Selene, powered by NVIDIA architecture, are not only breaking computational power records but, more importantly, opening doors to simulations at previously impossible scales. How do they work? What are they already being used for? And what challenges lie ahead for them in the coming years?
Since April 2022, when Frontier at the Oak Ridge National Laboratory became the first supercomputer to cross the 1 exaflops threshold, the world of high-performance computing (HPC) has entered a new era. This would not have been possible without NVIDIA architecture – from Hopper GPU chips to NVLink interconnects and AI-optimized software. But what do these achievements really mean for science and industry? And why is Europe, led by the JUPITER project, investing billions in exascale supercomputers?
NVIDIA supercomputers: an overview of key exascale systems
Before we dive into the applications and challenges, it is worth taking a look at the machines themselves. Here are the most important supercomputers based on NVIDIA technology that are already operational or will soon change the face of science:
1. JUPITER – the European giant from Jülich
The JUPITER (Jülich Universal Petascale Infrastructure for Advanced Research) project is not just another supercomputer – it is a cornerstone of the European exascale computing strategy. Being built in Germany by the Jülich Supercomputing Centre, it is scheduled to be launched in two phases:
- 2024: Partial launch with a performance of around 500 petaflops.
- 2025: Full power – approximately 1 exaflops.
What sets it apart? A hybrid architecture based on the NVIDIA Grace Hopper Superchip, which combines CPU and GPU in a single package. Thanks to this, JUPITER is expected to achieve not only massive computational power but also significantly lower energy consumption than traditional systems.
JUPITER will not be alone, however. It is part of the European eurohpc program, which aims to create several exascale-class supercomputers by 2025. The others include LUMI (Finland) and marenostrum 5 (Spain).
2. Frontier – the American power leader
Frontier, launched in May 2022 at the Oak Ridge National Laboratory, was the first supercomputer to officially exceed 1 exaflops (specifically 1.194 exaflops). Although its architecture relies on AMD EPYC processors and NVIDIA MI250X accelerators, it is NVIDIA's graphics chips that play a key role in its performance.
Frontier is not just a demonstration of power. It is already being used for:
- Modeling nuclear fusion (in collaboration with ITER).
- Simulating climate change with unprecedented resolution.
- Studying particles in high-energy physics (e.g., analyzing data from the LHC at CERN).
3. Selene – NVIDIA's laboratory
Selene is a supercomputer built by NVIDIA itself, primarily used for testing new technologies. Although its power is "only" 30–63 petaflops (depending on the configuration), thanks to A100 and A30 chips, it has become a platform for research into:
- New artificial intelligence algorithms.
- Optimization of scientific computing using Tensor Cores.
- Solutions in the field of precision medicine (e.g., simulations of drug interactions at the molecular level).
Selene was also used to develop the first molecular docking models in the fight against COVID-19, which significantly accelerated research into antiviral drugs.
4. EOS – a taste of zettascale
The latest addition to NVIDIA's arsenal is EOS – a system with a theoretical performance of 2.5 exaflops, launched in 2023. EOS is based on the NVIDIA Grace Hopper Superchip and is intended primarily for:
- Developing new AI models for physical simulations.
- Testing architecture for future zettascale supercomputers (10³¹ FLOPS).
- Accelerating calculations in fields such as astrophysics or synthetic biology.
NVIDIA architecture: why do HPC and AI go hand in hand?
The key to the success of exascale supercomputers is not just raw computational power, but also an architecture that enables efficient work. For years, NVIDIA has been developing technologies that combine high-performance computing (HPC) with artificial intelligence. Here is how it works:
1. Hopper GPU – the heart of the new generation
Introduced in 2022, the Hopper (GH200, Grace Hopper Superchip) chip is a combination of an NVIDIA Grace CPU and a Hopper GPU on a single chip. Its most important features are:
- Memory bandwidth: Up to 900 GB/s thanks to HBM3e memory, allowing for lightning-fast data transfer between CPU and GPU.
- 4th Gen Tensor Cores: Optimized for AI calculations, but also for scientific simulations (e.g., fluid dynamics, molecular modeling).
- NVLink-C2C: A new standard for CPU-to-GPU interconnects, offering 900 GB/s of bandwidth – crucial for the scalability of exascale systems.
Thanks to this, Hopper allows for 30% lower energy consumption at the same computational power compared to traditional solutions.
2. CUDA and the HPC ecosystem
The foundation of NVIDIA's software for supercomputers is CUDA – a GPU programming platform that enables the writing of efficient scientific applications. In 2023, CUDA 12.0 was introduced, which brings:
- Better optimization for floating-point calculations (FP64), crucial for physical simulations.
- Support for new AI algorithms, such as Diffusion Models or Physics-Informed Neural Networks (PINNs).
- NVIDIA HPC SDK – a suite of tools for compilers, debuggers, and mathematical libraries that are essential for writing efficient HPC code.
Additionally, NVIDIA offers NVIDIA Modulus – a framework for building AI models that can be used for physical simulations. An example is modeling turbulent flows in aviation, which can be trained on data from wind tunnel experiments.
3. Integrating AI with HPC
Traditional supercomputers were designed primarily for deterministic simulations (e.g., weather modeling). Today, thanks to AI, it is also possible to:
- Predict simulation results based on partial data (e.g., weather forecasting a week in advance in real-time).
- Generate synthetic data for model training (e.g., particle collision simulations in high-energy physics).
- Optimize energy consumption in data centers through predictive load management.
It is this synergy between HPC and AI that makes NVIDIA supercomputers so revolutionary. Artificial intelligence is becoming an inseparable element of scientific research, and NVIDIA provides the tools to use it effectively.
Science in the exascale era: projects that are already changing the world
Exascale supercomputers are not just a display of computational power – they are tools that are already revolutionizing scientific research. Here are a few key projects in which NVIDIA plays a major role:
1. Medicine: from COVID-19 to cancer drugs
The COVID-19 pandemic was the first major test for supercomputers in the fight against a global threat. Thanks to Selene and molecular docking algorithms, researchers were able to screen billions of drug combinations in just weeks – something that would have taken years using traditional methods. The results? Several promising antiviral drug candidates that underwent further clinical trials.
Currently, NVIDIA is collaborating with the National Cancer Institute in the USA on the Drug Discovery at Exascale project, which aims to:
- Simulate protein-drug interactions with unprecedented accuracy.
- Accelerate research into drugs for rare diseases and cancers.
Several promising molecules have already been identified and are currently being tested in laboratories.
2. Energy: nuclear fusion and renewable sources
Frontier is one of the key tools in research into nuclear fusion as part of the ITER project. Plasma simulations in a tokamak reactor require massive computational power because:
- Modeling plasma behavior in a magnetic field is a task requiring trillions of operations per second.
- AI is used to control the reaction and prevent instabilities that could lead to failure.
Thanks to supercomputers like Frontier, scientists are able to test hundreds of different plasma configurations virtually before moving to expensive physical experiments. This not only saves time but also millions of dollars.
Another important project is Wind farm optimization using JUPITER. Airflow simulations around wind turbines allow for:
- Reducing the noise generated by wind farms.
- Increasing energy efficiency through optimal turbine placement.
3. Climatology: next-generation weather models
Current weather models have a resolution of around 9 km, which makes it impossible to accurately predict local extreme events (e.g., storms or droughts). Thanks to JUPITER and similar systems, scientists will be able to create models with a resolution of 1 km – which will open doors to:
- Better forecasting of hurricanes and tornadoes a week in advance.
- Precise modeling of climate change at a regional level.
- Optimizing water management in the face of droughts.
The European Centre for Medium-Range Weather Forecasts (ECMWF) is collaborating with NVIDIA on the Destination Earth project, which aims to create a digital twin of the Earth – a virtual model of the planet that will allow for testing various climate scenarios.
4. Quantum physics and condensed matter
Exascale supercomputers also open new possibilities in research into quantum matter and superconductivity. An example is the Quantum Monte Carlo project on Frontier, which allows for:
- Simulating electron behavior in 2D materials (e.g., graphene).
- Studying high-temperature superconductors that could revolutionize energy.
Thanks to AI, scientists are able to predict material properties before they are synthesized in the laboratory – which significantly speeds up the discovery process.
Exascale challenges: energy, cooling, and software
Despite immense progress, exascale supercomputers face serious technological challenges. Here are those that NVIDIA and its partners are trying to solve:
1. Energy consumption: how to get below 50 MW?
Frontier consumes about 20 MW of energy, and planned zettascale systems may require as much as 50–100 MW. For comparison, that is as much as a medium-sized city! To address this, NVIDIA proposes:
- NVIDIA Grace Hopper Superchip: Up to 30% lower energy consumption than traditional CPU/GPU setups.
- Liquid immersion cooling: Submerging chips in a special fluid that dissipates heat 10 times more effectively than air.
- Use of renewable energy: Supercomputers like JUPITER will be powered primarily by wind and solar energy.
NVIDIA is also collaborating with energy companies on smart grids that will deliver energy at times of peak demand.
2. Cooling: how to handle a power density of 100 kW per rack?
In traditional data centers, power density is around 10 kW per rack. In exascale supercomputers, it is 100 kW or more – which creates huge cooling challenges. Solutions include:
- Two-phase cooling: Using boiling liquid that absorbs heat during evaporation.
- Direct-to-chip cooling: Fluid flowing directly through electronic circuits (tested in Selene).
- Waste heat utilization: Recovering thermal energy to heat buildings (e.g., in a data center in Luxembourg).
3. Software: how to write code that scales to exaflops?
Writing software for exascale supercomputers is one of the biggest challenges. Traditional languages like Fortran or C++ must be supplemented with:
- NVIDIA HPC SDK: A suite of tools for compiling, debugging, and optimizing code.
- openacc/openmp: Parallel programming standards that are compatible with NVIDIA GPUs.
- AI-Enhanced Simulation: Frameworks like NVIDIA Modulus, which combine physical simulations with machine learning models.
Another problem is porting existing applications to the GPU architecture. NVIDIA offers compilers that automatically convert CPU code to GPU, but this is not always optimal. That is why collaboration with scientists to adapt their algorithms to the new infrastructure is so important.
4. Costs: are exascale supercomputers economically justified?
Building an exascale supercomputer costs between $300 million and $1 billion. On top of that, there are maintenance costs ($20–50 million per year). Is it worth it? Yes, if it:
- Accelerates scientific discoveries that have a direct impact on the economy (e.g., new drugs, more efficient solar cells).
- Reduces the costs of physical experiments (e.g., testing aircraft prototypes virtually instead of in wind tunnels).
- Opens up new economic opportunities, such as the digital industrial revolution (e.g., smart cities, autonomous vehicles).
An example is Frontier, which paid for itself within the first two years thanks to:
- Savings in nuclear fusion research.
- Accelerated work on cancer drugs.
NVIDIA vs. the competition: who will win the exascale race?
NVIDIA is not the only player in the exascale supercomputer market. How does it compare to other technologies? Here is a comparison:
1. AMD – power at a lower cost
AMD, with EPYC chips and MI300X accelerators, is part of several exascale supercomputers, such as El Capitan in the USA (planned power: 2 exaflops). The advantages of AMD are:
- Higher computational density (more cores in the same space).
- Better price per FLOPS (which is crucial for government budgets).
The downsides? Less mature software and less AI support compared to NVIDIA.
2. Intel – delays and ambitious plans
Intel, with the Aurora project (2 exaflops, planned for 2024), is trying to compete with NVIDIA and AMD. Its advantage is:
- Integration with the x86 ecosystem, which makes porting existing applications easier.
- Ponte Vecchio chips (GPUs based on the Xe architecture).
The problem? Delivery delays and lower performance in AI calculations compared to NVIDIA Hopper.
3. Fujitsu – a leader in Japan, but not in exascale
Fugaku, a Fujitsu supercomputer based on A64FX chips, was the fastest supercomputer in the world in 2020–2021 (0.442 exaflops). However, its architecture does not scale well to the exascale level, which makes Fujitsu lose significance in the race.
4. Chinese supercomputers – mystery and ambition
China, which has been breaking records in Top500 for years, is building exascale supercomputers based on domestic chips. However, due to export restrictions and a lack of transparency, it is difficult to assess their actual performance and applications. They will most likely be used primarily for military and national security purposes.
In summary, NVIDIA remains the leader in the field of exascale supercomputers thanks to:
- Best performance in AI and HPC calculations.
- A mature software ecosystem (CUDA, HPC SDK).
- Strong partnerships with scientific and industrial institutions.
The future of exascale: what awaits us after 2025?
Exascale supercomputers are just the beginning. NVIDIA and its competitors are already working on the next revolutions:
1. Zettascale computing (10³¹ FLOPS)
By 2030, scientists plan to create the first supercomputers with zettascale power. To achieve this, the following will be necessary:
- New memory technologies (e.g., optical memory to replace HBM).
- Quantum integrated circuits for specialized calculations.
- Advanced AI algorithms that will be able to "understand" and predict physical phenomena on an unprecedented scale.
NVIDIA has announced the introduction of Blackwell chips in 2024, which are expected to be a milestone toward zettascale.
2. Supercomputers as a service (HPC-as-a-Service)
Currently, supercomputers are available primarily to scientific and government institutions. In the future, we can expect a pay-per-use model, where companies will be able to rent computational power on demand. An example is already NVIDIA DGX Cloud, which offers access to Hopper GPUs in the cloud.
3. Ethics and security of exascale
With the growing power of supercomputers, ethical challenges also arise:
- Military use: Supercomputers can be used to design nuclear weapons, cyberattacks, or mass surveillance systems.
- Control over AI: How to ensure that autonomous systems making scientific decisions do not operate outside of human control?
- Accessibility: How to avoid a situation where only a few countries and corporations have access to the most powerful research tools?
These questions will be crucial in the debate on the future of exascale.
Summary: why are NVIDIA supercomputers the future of science?
Exascale supercomputers based on NVIDIA technology are not just tools – they are a revolution in the approach to scientific research. Thanks to them, simulations that seemed like science fiction just a few years ago are becoming possible:
- Modeling the entire universe to understand dark matter.
- Rapidly discovering new drugs for previously incurable diseases.
- Optimizing the global energy economy through precise weather forecasts and designing more efficient technologies.
- Creating a virtual digital twin of the Earth that will help fight climate change.
Of course, the road to fully utilizing the potential of exascale is not easy. Challenges related to energy, cooling, software, and costs remain enormous. However, thanks to innovations like the Grace Hopper Superchip, NVLink-C2C, or NVIDIA Modulus, the world is on the verge of a new era of scientific discovery.
Will Europe, led by the JUPITER project, catch up with the United States and China? Will exascale supercomputers change our lives as much as the internet did? One thing is certain: the future of science is being written now – on NVIDIA GPUs.
Sources
- https://blogs.nvidia.com/blog/jupiter-exascale-supercomputing-science/
- https://www.olcf.ornl.gov/frontier/
- https://nvidianews.nvidia.com/news/nvidia-unveils-eos-2-5-exaflops-ai-supercomputer-for-research
- https://developer.nvidia.com/hpc-sdk
- https://www.exascaleproject.org/
- https://www.nvidia.com/en-us/data-center/grace-hopper-superchip/
Comments