AIDC
DC AI 热点

全部 AI 动态

8月26日2026-08-26
AWS HPC✦ 精选规则精选06:33

Part 1: Managing Large-Scale LLM Training with AWS ParallelCluster

This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. Introduction The Korean Government announced a national AI initiative to provide high-performance GPU infrastructure for Korea’s national AI research teams. AWS was selected as a supplier of GPU resources […]

8月5日2026-08-05
AWS HPC✦ 精选规则精选08:00

Resilient HPC and ML on AWS: Running Tightly Coupled Workloads on Spot Instances

This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. This post was contributed by Santosh Kumar, Bhagyaraju Kasina, Dr. Sandeep Sovani and Dr. Max Starr Researchers and engineering teams running High Performance Computing (HPC) jobs face a constant […]

6月26日2026-06-26
AWS HPC✦ 精选规则精选08:15

Transforming HPC Operations with Intelligent Workload Orchestration on AWS

This post was contributed by Manu Pillai, Gloria Macia and Natalia Jimenez, PhD Organizations running high-performance computing (HPC) workloads today operate largely as they have for decades: users manually specify required compute specifications for each of their jobs. Users spend valuable time analyzing workload requirements, selecting instance types, and troubleshooting infrastructure issues – time that […]

6月8日2026-06-08
AWS HPC✦ 精选规则精选21:51

Reducing costs by 50% while processing population-scale genomics with Mountpoint for Amazon S3 and AWS Batch

This post was contributed by Kambiz Shahim, Ankit Kalyani, and Chris Wright. Oxford Nanopore Technologies used Mountpoint for Amazon S3, AWS Batch, and Nextflow to build EPI2ME Cloud, a managed compute environment for processing human genomes at population-scale reliably and securely while reducing computational costs by 50%. EPI2ME Cloud forms Oxford Nanopore’s suite of local […]

6月3日2026-06-03
AWS HPC✦ 精选规则精选00:29

Monitoring AWS Parallel Computing Service

This post was contributed by Ronald Hudson and Nate Haynes High Performance Computing (HPC) on AWS demands precise monitoring, like the racing telemetry used by Formula 1 teams to deliver results. Like race engineers tracking car performance, AWS Parallel Computing Service (AWS PCS) administrators must monitor computing metrics in real-time. This vigilance is critical because […]