Description
Expanse is a Dell integrated compute cluster, with AMD Rome processors, NVIDIA V100 GPUs, interconnected with Mellanox HDR InfiniBand in a hybrid fat-tree topology. The GPU component of Expanse features 52 GPU nodes, each containing four NVIDIA V100s (32 GB SMX2), connected via NVLINK, and dual 20-core Intel Xeon 6248 CPUs. They feature 1.6TB of NVMe storage and 256GB of DRAM per node. There is HDR100 connectivity to each node. The system also features 12PB of Lustre based performance storage (140GB/s aggregate), and 7PB of Ceph based object storage.
RP Description
Expanse GPU is a compute cluster consisting of 52 nodes, each with four NVIDIA V100 (32 GB SXM2) GPUs linked by NVLink and dual 20-core Intel Xeon 6248 CPUs. It is particularly well suited for multi-GPU AI and machine-learning training, and is often used for deep-learning and other workloads that scale across GPUs.
Jobs Information
Expanse GPU is a compute cluster consisting of 52 nodes, each with four NVIDIA V100 (32 GB SXM2) GPUs linked by NVLink and dual 20-core Intel Xeon 6248 CPUs. It is particularly well suited for multi-GPU AI and machine-learning training, and is often used for deep-learning and other workloads that scale across GPUs.
GPU nodes are allocated as a separate resource. The GPU nodes can be accessed via either the gpu or the gpu-shared partitions.
When users request 1 GPU, in gpu-shared partition, by default they will also receive, 1 CPU, and 1G memory.
For more information and example job scripts see Expanse Running Jobs.