ACES

Description

ACES is a Dell cluster with a rich accelerator testbed consisting of Intel Max GPUs (Graphics Processing Units), Intel FPGAs (Field Programmable Gate Arrays), NVIDIA H100 and A30 GPUs, NEC Vector Engines, NextSilicon co-processors, and Graphcore IPUs (Intelligence Processing Units). Researchers must be based in the US and associated with a US academic research institution.

The ACES cluster consists of compute nodes using a mix of the following processors:
Intel Xeon 8468 Sapphire Rapids processors
Intel Xeon Ice Lake 8352Y processors
Intel Xeon Cascade Lake 8268 processors
AMD Epyc Rome 7742 processors

The compute nodes are interconnected with NVIDIA NDR200 connections for MPI and access to the Lustre storage. The Intel Optane SSDs and all accelerators (except the Graphcore IPUs and NEC Vector Engines) are accessed using Liqid's composable infrustructre via PCIe (Peripheral Component Interconnect express) Gen4 and Gen5 fabrics.

Resource ID
1659
Global Resource ID
aces.tamu.access-ci.org
Resource Type
Compute
Latest Status
production
Latest Status Begin
Project Affiliation
ACCESS
Organization Name
Texas A&M University
RP Description

ACES is a composable compute testbed consisting of Intel Sapphire Rapids CPU nodes that draw accelerators on demand from a shared pool over a Liqid PCIe fabric (NVIDIA H100 and A30 GPUs, Intel Data Center GPU Max, FPGAs, and NextSilicon coprocessors), plus fixed nodes carrying Graphcore IPUs, NEC Vector Engines, and an NVIDIA Grace-Hopper superchip. It is particularly well suited for work that compares or combines accelerators rather than running on one fixed configuration, and is often used for AI and ML training, benchmarking on emerging hardware, and porting codes to new devices. It includes a broad AI/ML and scientific software stack, with interactive applications delivered through Open OnDemand.

MFA Required
On
External Storage

Texas A&M HPRC offers storage beyond the ACES quotas: Google Drive and Microsoft OneDrive at 25 GB each, and HPRC Long Term Storage as a paid dedicated tier. The published terms are written for Texas A&M users, so confirm what you are eligible for if your only affiliation is an ACCESS allocation. Current options and rates are in [ACES Extra Storage Options].

What ACES itself keeps, and for how long, is set out in [ACES Key Policies].

File Transfer Text

Globus is the recommended route for large or many-file transfers: it retries and resumes on its own, and it runs against the ACES collection ACCESS TAMU ACES DTN rather than a login node.

For smaller copies, scp, sftp, rsync, rclone and gdown run on the ACES login node, login.aces.hprc.tamu.edu. ACES has no separate command-line data transfer node, and login nodes cut off long-running processes, so use Globus for anything large.

Setup instructions for every method are in [ACES File Transfer].

File Transfer Methods
Transfer Method
Globus
Data Transfer Node / Globus Collection
ACCESS TAMU ACES DTN
Transfer Method
scp/sftp
Data Transfer Node / Globus Collection
login.aces.hprc.tamu.edu
Transfer Method
rsync
Data Transfer Node / Globus Collection
login.aces.hprc.tamu.edu
Transfer Method
rclone
Data Transfer Node / Globus Collection
login.aces.hprc.tamu.edu
Transfer Method
gdown
Data Transfer Node / Globus Collection
login.aces.hprc.tamu.edu
Storage Text

ACES gives you three spaces: a small Home for code, scripts, and compiling; Scratch for the files an active job reads and writes; and Project space shared across your group.

Only Home is backed up. Scratch and Project are not, and neither is meant for long-term storage, so keep your own copies of anything you cannot regenerate.

Every tier also carries a file-count limit alongside its capacity limit, and the file-count limit is often what you hit first.

For paths and quotas, see [ACES File Systems].

Jobs Information

You can run jobs at different sizes and durations on ACES. The following lists the different queues that you can submit to, describing how many nodes you get, how long you can run, the type of resources you get, and the average wait time.

Jobs are scheduled by Slurm. Each job picks a partition with --partition, which sets the hardware it runs on, and must request a core count and a walltime. Compute nodes in the main pool provide 96 cores and 512 GB of memory. If you give no walltime, the partition's default applies; run scontrol show partition <partition> on a login node to see the default and the maximum for that partition. For node and architecture details see the [ACES Hardware] page; for job files and submission commands see the [ACES Batch System] guide.

The pvc partition's nodes carry 2 to 8 Intel GPU Max cards, depending on the node.

The Grace-Hopper node (gh01) is not a batch partition. You reach it by connecting directly over SSH from an ACES login node.

The nec partition holds one node with 8 NEC Vector Engine Type 20B-P cards. Request cards with --gres=gpu:ve:N (1 to 8), interactively with srun --partition nec ... or in a batch script with #SBATCH --partition=nec.

ACES uses a composable Liqid PCIe fabric, so accelerators (H100, A30, PVC, FPGA, NextSilicon) are attached on demand to a shared pool of Sapphire Rapids compute nodes rather than being fixed to a partition. A node can belong to several partitions at once, so the per-queue node counts overlap and are not additive; they total more than the 110-node compute pool by design. For a live view, run sinfo -s -p <partition> or gpuavail on a login node.

Storage Filesystems
Directory
Home
File System Path
/home/<userid>
Quota size
10
Quota inode amount
10000
Backup Policy
Backed up daily
Notes
Small scripts, config files, not for general use
Directory
Scratch
File System Path
/scratch/user/<userid>
Quota size
1000
Quota inode amount
250000
Backup Policy
Not backed up
Notes
Primary working directory for jobs, not for long-term storage. Data is also removed when quotas are exceeded.
Directory
Project
File System Path
/scratch/group/<projectid>
Quota size
5000
Quota inode amount
500000
Purge Policy
Purged 90 days after allocation expires
Backup Policy
Not backed up
Notes
Shared storage for group members
Queue Specifications
Queue Name
nec
Purpose
Memory-bandwidth-bound, vectorizable code such as sparse linear algebra, where throughput limits performance more than compute.
CPU Type
Intel Xeon 8268 (Cascade Lake)
GPU Type
NEC Vector Engine Type 20B-P
GPU Count
8
GPU vRAM
48
CPU Count
48
Node RAM
768
Queue Name
bittware
Purpose
Custom hardware acceleration on FPGAs, for algorithms that gain from a tailored datapath instead of a fixed CPU or GPU.
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
None
GPU Count
0
GPU vRAM
0
CPU Count
96
Node RAM
512
Queue Name
nextsilicon
Purpose
Speeding up existing HPC code on a coprocessor that profiles the run and reconfigures to it, with no code rewrite (by request).
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
None
GPU Count
0
GPU vRAM
0
CPU Count
96
Node RAM
512
Queue Name
gpu
Purpose
GPU-accelerated work such as deep-learning training and CUDA applications.
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
NVIDIA H100
GPU Count
2
GPU vRAM
80
CPU Count
96
Node RAM
512
Queue Name
memverge
Purpose
Large-memory jobs that use MemVerge Memory Machine to expand RAM with Intel Optane, for working sets that exceed a node's normal memory.
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
None
GPU Count
0
GPU vRAM
0
CPU Count
96
Node RAM
512
Queue Name
pvc
Purpose
AI/ML and HPC written with Intel's oneAPI/SYCL, or CUDA code ported to Intel GPUs.
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
Intel Data Center GPU Max 1100 (Ponte Vecchio)
GPU Count
8
GPU vRAM
48
CPU Count
96
Node RAM
512
Queue Name
cpu
Purpose
CPU-only jobs that need no accelerator; the default choice for general compute.
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
None
GPU Count
0
GPU vRAM
0
CPU Count
96
Node RAM
512
Queue Name
gpu_debug
Purpose
Short GPU test and debug runs.
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
NVIDIA A30
GPU Count
2
GPU vRAM
24
CPU Count
96
Node RAM
512
Queue Name
gpu_debug
Purpose
Short GPU test and debug runs.
CPU Type
Intel Xeon 8468 (Sapphire Rapids)
GPU Type
NVIDIA H100
GPU Count
2
GPU vRAM
80
CPU Count
96
Node RAM
512
Queue Name
gh01
Purpose
Not a batch partition. GPU work that outgrows normal GPU memory, using Grace-Hopper's unified CPU-GPU memory and fast NVLink; reached over SSH, not Slurm.
CPU Type
ARM Neoverse V2 (NVIDIA Grace ARM)
GPU Type
NVIDIA H100
GPU Count
1
GPU vRAM
96
CPU Count
72
Node RAM
480
Datasets
Dataset Name
pytorch-computer-vision-datasets
Dataset Description

A collection of standard computer vision datasets formatted for PyTorch, supporting tasks like image classification and object detection. On ACES, these are used to benchmark GPU performance and test distributed deep learning workflows across accelerators.

Dataset Name
pytorch-language-modelling-datasets
Dataset Description

Text-based datasets for training NLP and language models in PyTorch. In ACES, they support benchmarking of large-scale, memory-intensive workloads and evaluating performance of transformer-based models across hardware.

Dataset Name
tensorflow-computer-vision-datasets
Dataset Description

Computer vision datasets optimized for TensorFlow, covering tasks such as classification and segmentation. Within ACES, they enable framework comparisons and validation of TensorFlow pipelines on heterogeneous accelerators.

Dataset Name
tensorflow-language-modelling-datasets
Dataset Description

NLP datasets prepared for TensorFlow, used for language modeling, translation, and text analysis. On ACES, they help evaluate distributed training performance and accelerator efficiency for sequential data workloads.

Dataset Name
videollama_dataset
Dataset Description

A multimodal dataset combining video and text for tasks like video understanding and captioning. In ACES, it is used to test high-throughput, multi-accelerator workflows and benchmark complex AI pipelines.