1 to 25 of 28 Slurm Workload Manager Jobs in the UK

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Senior Software Engineer - Scientific / HPC

Hiring Organisation
Technical Futures Ltd
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 - £90,000 per annum
code quality. Some/most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Senior Machine Learning Systems Engineer (Frameworks & Tooling)

Location
Greater London, England, United Kingdom
allowance A monthly quality time allowance A track record of building tools that increase developer velocity for ML teamsExperience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JAX, AAAI, Nature, COLING, ACL, EMNLP)Experience ...

Infrastructure Engineer

Location
Cambridge, England, United Kingdom
LDAP, NIS and DHCP. Experience with NFS/enterprise storage. FlexLM experience is desirable. Experience with Ansible, Puppet, Chef or Salt. Knowledge of Slurm, Grid Engine, LSF or similar is advantageous. #J-18808-Ljbffr ...

Infrastructure Engineer

Location
Cambridge, England, United Kingdom
/reliable systems, Care about the experience of the people relying on your infrastructure. Nice to have/Beneficial HPC‐style job scheduling (Kueue, Slurm, MPI, or similar). Experience with large scientific or array datasets. Any background in photonics, semiconductors, EDA, or scientific computing. We are particularly interested ...

HPC & AI Platform Engineer – GPU Clusters

Location
United Kingdom
ensure seamless AI platform operations. You will design, deploy, and manage large-scale HPC and GPU-accelerated clusters (NVIDIA). You will implement Slurm-based scheduling, optimize InfiniBand/Ethernet networks, and lead automation, monitoring, and incident response across multi-vendor stacks #J-18808-Ljbffr ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. Work with Slurm, Kubernetes, NVIDIA GPU Operator, NCCL and GPUDirect. Design high-performance storage solutions for AI workloads. Lead technical discussions with CTOs, AI leaders and infrastructure ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

HPC & AI Platform Engineer (GPU/Networking)

Location
United Kingdom
with vendor engineering teams to ensure seamless AI platform operations. You will design, deploy and manage large-scale GPU-accelerated clusters using NVIDIA GPUs, Slurm, InfiniBand and high-availability practices, while automating provisioning, monitoring, and security. #J-18808-Ljbffr ...

MLOps Engineer

Location
Oxford, England, United Kingdom
across on-premises accelerator clusters and cloud (GPU/CPU) for training and simulation workloads Drive infrastructure-as-code practices: containerisation, orchestration (Kubernetes/Slurm), and reproducible environment management Contribute to the internal developer platform: self-service tooling, documentation, and runbooks that raise engineering productivity across the company What … Experience with experiment tracking and model lifecycle management tools (MLflow, W&B, DVC, or similar) Solid understanding of containerisation (Docker) and orchestration (Kubernetes or Slurm) for distributed compute workloads Infrastructure-as-code mindset: Terraform, Ansible, or equivalent; CI/CD pipelines (GitHub Actions, Jenkins, or similar) Experience with hardware ...

Platform Engineer

Location
United Kingdom
managing large‐scale HPC and GPU‐accelerated clusters, including NVIDIA based compute environments. Implementing and administering HPC scheduling and resource‐management systems (e.g., Slurm), including GPU partitioning, workload scheduling, and capacity planning. Architecting and optimising InfiniBand and Ethernet network topologies. Ensuring high availability and resilience through failover strategies ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Associates (UK) Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£700 - £765 per day
data-centre teams. Key Requirements Strong Platform/Infrastructure Architecture experience across compute, storage, networking and Linux. Expert-level Kubernetes architecture experience. Strong Slurm and Run experience - essential. Proven experience with GPU/HPC environments and large-scale AI platforms. Hands-on experience with NVIDIA HGX GB300/NVL72 … including NVLink, NVSwitch and Grace Blackwell architecture. Experience with NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Technology
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£700.00 - £765.00 per day
data-centre teams. Key Requirements Strong Platform/Infrastructure Architecture experience across compute, storage, networking and Linux. Expert-level Kubernetes architecture experience. Strong Slurm and Run experience - essential. Proven experience with GPU/HPC environments and large-scale AI platforms. Hands-on experience with NVIDIA HGX GB300/NVL72 … including NVLink, NVSwitch and Grace Blackwell architecture. Experience with NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform ...

R&D Solution Architect

Location
Hook, England, United Kingdom
implementing solutions that adhere to FAIR data principles (Findable, Accessible, Interoperable, Reusable)." - Experience architecting for High-Performance Computing (HPC) environments, including knowledge of workload schedulers (e.g., SLURM) and applying cloud-native patterns to scientific, batch-processing workloads. - Familiarity with scientific workflow management tools (e.g., Nextflow, Snakemake ...

Senior Solution Engineer – GPU & AI Infrastructure

Location
United Kingdom
role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare …/Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing). Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native/Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow). Storage Integration: Architect high-bandwidth parallel ...

Senior HPC Engineer - Hybrid - Inside IR35

Hiring Organisation
Hamilton Barnes
Location
Stevenage, Hertfordshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP 340 Daily
Role We are looking for an experienced Senior HPC Engineer to support and maintain a scientific computing environment, with a strong focus on RHEL, Slurm and HPC infrastructure . You will work closely with research scientists and technical teams to ensure HPC services are secure, reliable and high performing. … Responsibilities Administer, patch and maintain RHEL 7, 8 and 9 across HPC clusters and workstations. Deploy, configure and manage Slurm , including queues, partitions and scheduling. Monitor cluster health, performance, storage, networking and resource utilisation. Install and support scientific applications, compilers, libraries and MPI environments. Work with scientists to optimise ...

Solutions Engineer

Location
Greater London, England, United Kingdom
engineering, and operations teams to ensure proposed solutions are realistic, scalable, and aligned with platform standards Provide guidance on compute, networking, storage, orchestration, and workload optimisation for AI and machine learning use cases Help create repeatable demo environments, technical playbooks, reference architectures, and sales enablement materials … data centre, or platform environments Good understanding of cloud infrastructure, GPU compute, AI/ML workloads, or high-performance infrastructure Familiarity with containers, Kubernetes, Slurm, orchestration platforms, or workload deployment models Understanding of networking, storage, and distributed compute concepts in modern infrastructure environments Ability to quickly learn ...

Senior Staff+ Software Engineer, Kubernetes Platform

Location
Greater London, England, United Kingdom
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us Design, build, and operate core cluster services such ...

Linux/RHEL Engineer - HPC

Hiring Organisation
Gazelle Global Consulting Ltd
Location
Hertfordshire, South East, United Kingdom
Employment Type
Contract
Senior Linux/HPC Engineer to support a scientific computing environment in Stevenage. This hands-on role covers Linux infrastructure, HPC clusters, workload scheduling and scientific application support, working closely with research and technical teams. You'll keep the environment secure, stable and performing effectively, while troubleshooting infrastructure … application issues and improving cluster utilisation and job throughput. Key Skills: Strong RHEL 7, 8 and 9 administration and troubleshooting HPC cluster management and Slurm configuration Scientific or research application support in Linux HPC environments MPI, compilers and scientific libraries Hardware, OS, scheduler and application troubleshooting ServiceNow or equivalent ...

Senior HPC Engineer

Hiring Organisation
Gazelle Global Consulting Ltd
Location
Stevenage, Hertfordshire, South East, United Kingdom
Employment Type
Contract, Work From Home
supporting scientific applications and complex computational workloads. Key Skills: Strong hands-on administration of RHEL 7, 8 and 9 HPC cluster administration and troubleshooting Slurm deployment, configuration and workload management Scientific applications within Linux HPC environments MPI libraries and computational workloads Hardware, OS and scheduler troubleshooting ServiceNow ...

Research Software Engineer

Location
Greater London, England, United Kingdom
surface live metrics. Write efficient, well-tested Python and systems code; enforce code review, CI, and observability. Design and optimise distributed services (Kubernetes/SLURM, thousands-of-GPU jobs). Prototype utilities (CLI, dashboards) and carry them through to stable, shared libraries. About the Research Engineering team Based … observability. Fluency in Python plus one systems language (C++, Rust, Go or Java). Hands-on with container orchestration and schedulers (Kubernetes/K8s, SLURM, or similar). Comfortable profiling performance, optimising I/O, and automating workflows. Self-starter, low-ego, collaborative, high-energy. Nice-to-haves Exposure ...

Senior Software Engineer - Research Technology

Location
Greater London, England, United Kingdom
fundamentals: data structures, algorithms, networking, OS, concurrency, and system design. Experience running compute at cluster scale: job scheduling, resource management, retries, and reliability. Slurm, Kubernetes, Ray, Spark, or custom internal schedulers all count. Proven data-engineering experience: schema design, storage formats, compression, I/O trade-offs, and pipelines … ship production software safely and repeatedly, with an obsession for data driven quality. Desirable/nice-to-have Rust experience alongside C++ and Python. Slurm or other cluster scheduler expertise. Familiarity with ML/Deep Learning frameworks. Prior finance or market-data experience, including low-level market connectivity. ...