1 to 25 of 76 Slurm Workload Manager Jobs in the UK

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Machine Learning Engineer (Mid to Principal)

Location
West of England, England, United Kingdom
Jira. Edge deployments: Nvidia Jetson (e.g. AGX Orin), Raspberry Pi, or other embedded accelerators. Distributed model training & infra: Pytorch DDP, FDSP and TorchTitan, Megatron, Slurm, Run:ai, DeepSpeed, Kubernetes, cloud or on‐prem GPU clusters. About you You’ve built ML systems that persist—deployed in real settings, iterated ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub. Knowledge of other ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

HPC System Administrator

Location
Farnborough, England, United Kingdom
defects. Perform Software and firmware testing for any fixes, upgrades, security patch. Position requirements: Hands-on experience with HPC tools and schedulers such as Slurm and PBS Pro, as well as Confluent Solid scripting and automation skills using Ansible, Bash, and Python Experience with distributed storage systems and file ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
environments Enjoys solving complex technical problems collaboratively Nice-to-Haves Experience with large scale model training or distributed inference systems Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms Experience with InfiniBand, RDMA, or high-performance networking Experience operating bare metal infrastructure Familiarity with storage systems commonly used ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms. ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron … continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Senior HPC Systems Administrator Closing date: Oct 01, 2026

Location
Oxford, England, United Kingdom
systems integration skills Excellent written and verbal communication abilities Ability to work independently and collaboratively in a research environment Desirable Skills Experience with SLURM or other job scheduling systems Experience with containerisation and virtualisation technologies Knowledge of cloud computing platforms Familiarity with scientific computing tools and AI/ ...

Senior HPC Systems Administrator

Location
Oxford, England, United Kingdom
Strong troubleshooting and systems integration skills Excellent written and verbal communication abilities Ability to work independently and collaboratively in a research environment Experience with SLURM or other job scheduling systems Experience with containerisation and virtualisation technologies Knowledge of cloud computing platforms Familiarity with scientific computing tools and AI/ ...

Senior HPC Linux Engineer — On‐Site, Security‐Clearance Ready

Location
West of England, England, United Kingdom
system security and performance monitoring. The role requires a 3+ year Linux background, scripting skills (Bash/Python), and a willingness to learn Slurm, xCAT and Ansible. On-site work five days a week with UK National status and clearance eligibility. #J-18808-Ljbffr ...

Senior Engineer, Systematic Data

Hiring Organisation
Balyasny Asset Management
Location
London, UK
Employment Type
Full-time
programming languages (e.g. Java/C#) Experience with datasets used for cash equity research and trading Experience with distributed computing (e.g. Spark/Dask, Slurm/HTCondor) Experience with event-driven, asynchronous architectures and related messaging technologies (e.g. Kafka, MQ) Experience with cloud platforms (e.g. AWS, GCP, Azure) Experience ...

HPC Systems Engineer

Location
Greater London, England, United Kingdom
Need: At least 5+ years of professional experience in high performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required At least 5+ years of experience with Linux systems administration High proficiency ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, UK
Employment Type
Full-time
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity ...

Hybrid HPC System Administrator – Slurm, Linux & Storage

Location
Farnborough, England, United Kingdom
Lenovo in the United Kingdom (Hampshire) is hiring an HPC System Administrator to manage data center infrastructure, Linux/Unix environments, and monitoring to ensure uptime and security. This role blends on-site and remote ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation. Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential. Relevant Background: Experience ...

Research Engineer, Forge

Location
Greater London, England, United Kingdom
comfort in fast‐moving, under‐specified environments. Nice to have Distributed training experience (FSDP, DeepSpeed, Megatron, etc.). Cluster/orchestration experience (SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, etc.). Experience building reliable ML infrastructure, evaluation systems, or large‐scale data processing pipelines. Research experience in LLMs, agents, multimodal ...

Senior Data & MLOps Engineer

Location
Greater London, England, United Kingdom
least one systems language (Go, Rust, or C++). Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm). Familiarity with GPU telemetry systems such as NVML or DCGM and hardware‐level monitoring concepts. Demonstrated experience scaling systems from Proof‐of‐Concept ...

Senior Field Engineer - AI/ML HPC & Kubernetes Solutions

Location
Greater London, England, United Kingdom
drive proofs of concept to accelerate client deployments. You will collaborate with engineering teams, contribute to product direction, and help optimize workloads using Slurm, NCCL, and Infiniband in high-performance environments. #J-18808-Ljbffr ...

AI infrastructure engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
throughput, and ensuring training is as efficient and cost-effective as possible. You'll also play a critical role in managing cluster orchestration with Slurm and Kubernetes while helping evolve the platform to support next-generation GPU infrastructure and specialised compute providers. This is an opportunity to work across ...

AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
will eliminate bottlenecks in the data path to ensure training is fast and as capital efficient as possible alongside managing cluster orchestration using slurm and Kubernetes while preparing to expand into specialised GPU providers. And finally you will master the stack from pytorch based learning libraries to complex data ...

Senior Software Engineer Scientific Data Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
Experience: Experience leveraging Kubernetes for application deployment, and familiarity with distributed computing frameworks (e.g., Ray, Spark), or specialised batch schedulers/resource managers (e.g., Slurm, Volcano, Kueue). Demonstratedexpertisewith specialised serving engines (e.g.,vLLM,SGLang) or techniques for deploying models in resource-constrained or high-throughput environments. Familiarity withGitOpsand ...

Bioinformatician Plant Biology Institute Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
bioinformatics tools for genome assembly, annotation and variant analysis. Familiarity operating within Linux environments, including operating within HPC and/or cloud environments (e.g. SLURM, OCI). Experience applying machine learning methods to genomic and transcriptomic data, with an understanding of different ML approaches and their suitability for different ...

Infrastructure Engineer

Location
Cambridge, England, United Kingdom
/reliable systems, Care about the experience of the people relying on your infrastructure. Nice to have/Beneficial HPC‐style job scheduling (Kueue, Slurm, MPI, or similar). Experience with large scientific or array datasets. Any background in photonics, semiconductors, EDA, or scientific computing. We are particularly interested ...