1 to 25 of 51 Slurm Workload Manager Jobs in London

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 70 K
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
environments Enjoys solving complex technical problems collaboratively Nice-to-Haves Experience with large scale model training or distributed inference systems Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms Experience with InfiniBand, RDMA, or high-performance networking Experience operating bare metal infrastructure Familiarity with storage systems commonly used ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads.AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms.ML frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent … ArgoCD) for continuous, declarative platform delivery.Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines.Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms. ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron … continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Research HPC Support Engineer

Hiring Organisation
The London School of Economics and Political Science (LSE)
Location
London, United Kingdom
Salary
£ 70 K
have:- Strong technical expertise in high-performance computing (HPC) and GPU systems. - Proven experience administering, configuring, and optimising HPC clusters and GPU systems (e.g. Slurm, OpenPBS, K8).- Hands-on experience with open-source software build and installation frameworks specifically designed for High-Performance Computing (HPC) environments, such ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, UK
Employment Type
Full-time
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation. Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential. Relevant Background: Experience ...

Research Engineer, Forge

Location
Greater London, England, United Kingdom
comfort in fast‐moving, under‐specified environments. Nice to have Distributed training experience (FSDP, DeepSpeed, Megatron, etc.). Cluster/orchestration experience (SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, etc.). Experience building reliable ML infrastructure, evaluation systems, or large‐scale data processing pipelines. Research experience in LLMs, agents, multimodal ...

Senior Data & MLOps Engineer

Location
Greater London, England, United Kingdom
least one systems language (Go, Rust, or C++). Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm). Familiarity with GPU telemetry systems such as NVML or DCGM and hardware‐level monitoring concepts. Demonstrated experience scaling systems from Proof‐of‐Concept ...

HPC Systems Engineer: Research Compute & AI Workloads

Location
Greater London, England, United Kingdom
support a large scale computing environment used for data-intensive research, modelling and AI workloads. You will work across Linux compute environments, GPU infrastructure, Slurm, and containerisation with Docker or Singularity, driving automation and platform improvements to better serve researchers. #J-18808-Ljbffr ...

Senior Field Engineer - AI/ML HPC & Kubernetes Solutions

Location
Greater London, England, United Kingdom
drive proofs of concept to accelerate client deployments. You will collaborate with engineering teams, contribute to product direction, and help optimize workloads using Slurm, NCCL, and Infiniband in high-performance environments. #J-18808-Ljbffr ...

AI infrastructure engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
throughput, and ensuring training is as efficient and cost-effective as possible. You'll also play a critical role in managing cluster orchestration with Slurm and Kubernetes while helping evolve the platform to support next-generation GPU infrastructure and specialised compute providers. This is an opportunity to work across ...

Member of Technical Staff - Research Software Engineer

Hiring Organisation
Reflection
Location
London, UK
Employment Type
Full-time
pipeline, expert) Large-scale distributed training infrastructure Communication optimization (NCCL, RDMA, GPU interconnects) FSDP/ZeRO and model sharding Orchestration & Runtime Systems Ray, Kubernetes, Slurm Distributed runtimes and async systems Containerization and sandboxing Frameworks PyTorch JAX Megatron-style training stacks Triton/custom kernels Data Infrastructure Large-scale dataset ...

IT Lead Engineer London

Location
Greater London, England, United Kingdom
slow" needs a real root cause, not a restart. What you will do Own and evolve our HPC environment: cluster administration, job scheduling (e.g., Slurm/PBS/LSF), performance tuning, and capacity planning for compute-heavy engineering workloads. Experience with Entra ID Governance: Access Reviews, Identity Protection, Privileged ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, UK
Employment Type
Full-time
cost-efficiency. Your ImpactProvision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed networking … from individuals. Anaplan does not: Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person. Send job offers via email. All offers are first extended verbally by a member of our internal recruitment team ...

Senior HPC Research Engineer: Accelerate Science

Location
Greater London, England, United Kingdom
reproducible computational solutions and maintain resilient services across bioinformatics, AI/ML, data processing, and simulations. The role requires hands‐on Linux HPC administration, SLURM knowledge, and experience with #J-18808-Ljbffr ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused computeBenefitsCompetitive salary and eligibility for equityBiannual bonus schemeFully expensed tech to match your needsPrivate health ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute Competitive salary and eligibility for equity Biannual bonus scheme Fully expensed tech to match ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

Senior Lead Software Engineer (C++, Python) – HPC & Cloud

Location
Greater London, England, United Kingdom
GenAI to improve monitoring, management, and efficiency of large compute platforms. You will collaborate with cross-functional teams, deploy on Kubernetes, and work with Slurm, LSF, Spark, Ray, and Symphony in cloud environments to deliver scalable, secure technologies #J-18808-Ljbffr ...

Senior HPC Engineer for Biomedical Computation

Hiring Organisation
MRC Laboratory of Medical Sciences
Location
London, UK
Employment Type
Full-time
invites applications for a Senior Research HPC Engineer to design, manage, and develop LMS's next‐generation HPC infrastructure. You will work across Linux, SLURM, and cluster automation to improve performance, reliability, and accessibility for biomedical researchers. The role combines hands‐on systems engineering with collaboration with scientists ...

Senior Research HPC Engineer - Accelerate Discovery

Location
Greater London, England, United Kingdom
design, deploy, and maintain HPC services, containers, and workflows, aligning with strategic goals and research needs. The role requires extensive experience with Linux HPC, SLURM, and software packaging, plus strong collaboration with researchers and cross‐functional IT teams to deliver reliable, high‐performance solutions. #J-18808-Ljbffr ...

Senior Research HPC Engineer

Hiring Organisation
MRC Laboratory of Medical Sciences
Location
London, United Kingdom
Employment Type
Permanent
Salary
£65,000
operates its own dedicated HPC environment, which is extensively used by multidisciplinary research groups and Institute facilities. The LMS IT department currently manages a SLURM-based cluster comprising CPU, high-memory (HMEM), and GPU nodes, supporting a diverse range of computational research. About the role As the Institutes computational ...