Arcadion

Pharmaceutical & Life Sciences

Secure Computing Power for Scientific Research

Deploying on-premises HPC and GPU infrastructure for research teams.

Long aisle of illuminated server racks in a data centre
Organization
Jubilant Pharma
Industry
Pharmaceutical & Life Sciences
Project Focus
HPC, GPU Infrastructure & Scientific Workload Management
Environment
On-Premises Linux HPC & GPU Server

The Challenge

Hardware Was Only Part of the Requirement

Scientific research teams increasingly rely on compute-intensive applications for data analysis, machine learning, modelling and GPU-accelerated experimentation. Jubilant Pharma required a dedicated high-performance computing environment that could support these workloads while providing fair, controlled access to shared resources. Installing the hardware was only one part of the requirement. The organization also needed:

A production-ready Linux operating environment
Centralized scheduling for CPU, memory and GPU workloads
Secure remote access for researchers and administrators
Controlled allocation of shared GPU resources
Repeatable software environments for scientific applications
Operational documentation and practical training for the science team
Ongoing monitoring and support for the hardware, OS and scheduler

Without a workload scheduler, individual researchers could consume large portions of the server's CPU, memory or GPU capacity, affecting other users and making resource utilization difficult to manage.

The Action

From Hardware Validation to Managed Operations

Arcadion deployed and operationalized a dedicated Dell PowerEdge R760XA HPC server, creating a secure on-premises computing platform for the client's scientific research workloads. The environment was built around Linux, NVIDIA GPU acceleration and Slurm workload scheduling. Slurm was configured to control how researchers requested CPU, memory, runtime and GPU resources, helping the organization provide predictable and equitable access to its shared computing infrastructure.

1 Dedicated HPC & GPU compute node
192 Logical CPU resources
503 GiB Schedulable system memory
14 NVIDIA MIG GPU slices from 2 physical GPUs
Slurm for workload scheduling and resource allocation
Lmod for controlled software and application environments
Secure SSH-based user access
Dedicated storage for user files and projects
Temporary and scratch data structures
Structured job output locations

Arcadion began by validating the physical server, firmware, BIOS configuration and core hardware health.

Secure remote management and VPN access were established so the HPC engineering team could configure and support the environment without exposing the management interfaces directly to the internet.

This provided a stable and secure foundation before the operating system and workload management services were introduced.

Engineer at a laptop showing a firmware update progress indicator

A Linux operating system was installed and configured specifically for high-performance computing workloads. The implementation included:

Server hostname and node identity configuration
SSH key-based authentication
Network Time Protocol synchronization
Kernel tuning and performance profiles
Internal network and node communication settings
Local, shared, home and scratch filesystem planning
Driver and OS preparation for GPU computing

Rather than deploying a standard Linux server, the environment was tuned around the performance, access and storage patterns expected from scientific computing.

Slurm was implemented as the central workload manager for the HPC platform. Researchers could submit interactive sessions, batch jobs and longer-running computational workloads while formally requesting the resources required by each job. This included:

CPU cores
System memory
Maximum runtime
GPU or MIG resources
Software environments
Output and logging locations

Slurm provided centralized job queuing, scheduling, resource isolation and workload visibility. It also reduced the risk of one user unintentionally consuming the entire server.

The two physical GPUs were configured using NVIDIA Multi-Instance GPU technology, creating 14 independently schedulable GPU slices.

Slurm Generic Resources, or GRES, was configured so researchers could request one or more GPU slices based on the needs of their experiment.

This allowed the organization to support several concurrent GPU workloads instead of dedicating an entire physical GPU to every researcher or job. It also improved resource utilization for experiments that did not require the full capacity of a GPU.

Macro detail of a GPU heatsink and server internals

Lmod environment modules were introduced to help the research team load and switch between software versions without modifying the server's base operating system.

Researchers could select approved versions of tools such as Python and CUDA for each workload. This helped maintain separation between application dependencies and improved the repeatability of scientific experiments.

Module configurations could also be included directly in Slurm batch scripts, ensuring that workloads started with the correct software environment regardless of who submitted the job.

A practical training program was delivered to help the science team adopt the new platform. Training covered the complete researcher workflow:

Connecting securely through SSH
Checking cluster and queue availability
Requesting CPU, memory and GPU resources
Starting interactive Slurm sessions
Submitting production workloads through sbatch
Loading Python, CUDA and other software modules
Monitoring GPU utilization
Reviewing job output and troubleshooting failures
Releasing resources after a job was completed

The team was also trained to use tmux, allowing researchers to preserve terminal sessions and computational work if their local SSH connection was interrupted.

A detailed 24-page Slurm Operations and Training Manual was produced for the environment. The runbook included:

An overview of the HPC architecture
Login access versus scheduled compute guidance
Core Slurm command references
GPU partition, GRES and MIG instructions
Interactive session examples
tmux operational guidance
Lmod software environment management
Production sbatch script templates
GPU monitoring procedures
Troubleshooting guidance
Resource request cheat sheets
Quick-reference cards for researchers and operators

The documentation was designed to support both first-time users and experienced researchers who needed a day-to-day operational reference.

Following deployment, Arcadion established ongoing operational support for the platform. The managed service included:

Hardware health monitoring
Power, temperature, disk and component monitoring
Firmware lifecycle planning
Linux security patching
Kernel and GPU driver maintenance
SSH and access control administration
Filesystem capacity monitoring
Slurm scheduler health monitoring
Queue and partition administration
User and project quota management
Priority and fair-share tuning
Job failure analysis and log review

Results and Business Impact

A Research Platform the Science Team Could Actually Use

Slurm gave Jubilant Pharma a structured method for allocating CPU, memory and GPU resources across its science team. Researchers could request the capacity they needed without manually coordinating access to the server.

Dividing the physical GPUs into 14 MIG slices allowed multiple researchers and experiments to use GPU resources concurrently. This improved the practical value of the hardware and reduced unnecessary resource contention.

Standardized software modules and batch job templates helped researchers define the software, compute resources and runtime requirements for each experiment. This improved consistency between runs and provided clearer operational records for troubleshooting and future research.

Memory limits, runtime limits and controlled GPU allocations reduced the risk of a single runaway process consuming the entire HPC server. The team also gained better visibility into active, pending and completed workloads.

Practical training, reusable scripts and a detailed operations manual enabled researchers to begin using the environment without becoming Linux or HPC infrastructure specialists.

Arcadion's managed HPC service provided the client with ongoing oversight of the server, Linux platform, Slurm scheduler and underlying hardware. This allowed the science team to remain focused on research instead of infrastructure administration.

Why Arcadion?

More than an HPC server installation. A complete scientific computing service.

Arcadion combines enterprise infrastructure experience with specialized knowledge in high-performance computing, GPU platforms, Linux engineering and managed technology services.

  • Secure infrastructure deployment
  • Linux and GPU platform configuration
  • Slurm workload scheduling
  • NVIDIA MIG resource management
  • Scientific software environment controls
  • Researcher training
  • Operational documentation
  • Ongoing managed support
Arcadion Chip Image

Ready to Power Your Scientific Research?

Let's build a secure, high-performance computing environment your team can rely on.

Connect