1 to 25 of 88 Slurm Workload Manager Jobs in the UK

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Machine Learning Engineer (Mid to Principal)

Location
West of England, England, United Kingdom
Jira. Edge deployments: Nvidia Jetson (e.g. AGX Orin), Raspberry Pi, or other embedded accelerators. Distributed model training & infra: Pytorch DDP, FDSP and TorchTitan, Megatron, Slurm, Run:ai, DeepSpeed, Kubernetes, cloud or on‐prem GPU clusters. About you You’ve built ML systems that persist—deployed in real settings, iterated ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 100 K
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub.Knowledge of other programming ...

Senior HPC/linux Engineer -

Hiring Organisation
Infoplus Technologies UK Ltd
Location
Stevenage, Hertfordshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP Annual
experienced Senior HPC Engineer to manage, secure and maintain the organisation's scientific computing environment. The role will focus on RHEL-based HPC infrastructure, Slurm workload management and scientific application support, while working closely with research scientists and technical teams to deliver reliable, secure and high-performing computing … Your responsibilities: Administer, patch, secure and maintain RHEL 7, 8 and 9 environments across HPC clusters and high-end workstations. Deploy, configure and manage Slurm, including scheduling, queues, partitions and fair-share policies. Monitor and optimise cluster health, resource utilisation, storage, networking and job throughput. Install and support scientific ...

HPC System Administrator

Location
Farnborough, England, United Kingdom
defects. Perform Software and firmware testing for any fixes, upgrades, security patch. Position requirements: Hands-on experience with HPC tools and schedulers such as Slurm and PBS Pro, as well as Confluent Solid scripting and automation skills using Ansible, Bash, and Python Experience with distributed storage systems and file ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 70 K
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads.AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms.ML frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent … ArgoCD) for continuous, declarative platform delivery.Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines.Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms. ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron … continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Senior HPC Linux Engineer — On‐Site, Security‐Clearance Ready

Location
West of England, England, United Kingdom
system security and performance monitoring. The role requires a 3+ year Linux background, scripting skills (Bash/Python), and a willingness to learn Slurm, xCAT and Ansible. On-site work five days a week with UK National status and clearance eligibility. #J-18808-Ljbffr ...

Senior Software Engineer - Scientific / HPC

Hiring Organisation
Technical Futures Ltd
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 - £90,000 per annum
code quality. Some/most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible ...

Research HPC Support Engineer

Hiring Organisation
The London School of Economics and Political Science (LSE)
Location
London, United Kingdom
Salary
£ 70 K
have:- Strong technical expertise in high-performance computing (HPC) and GPU systems. - Proven experience administering, configuring, and optimising HPC clusters and GPU systems (e.g. Slurm, OpenPBS, K8).- Hands-on experience with open-source software build and installation frameworks specifically designed for High-Performance Computing (HPC) environments, such ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
titles matter less than experience. We're particularly interested in people with experience in:Large-scale distributed training.PyTorch and modern deep learning frameworks.Kubernetes, Slurm or GPU orchestration platforms.AWS and specialist GPU cloud providers.High-performance computing and distributed systems.Training optimisation, memory management and networking.MLOps tooling, CI/CD and experimentation ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
matter less than experience. We're particularly interested in people with experience in: Large-scale distributed training. PyTorch and modern deep learning frameworks. Kubernetes, Slurm or GPU orchestration platforms. AWS and specialist GPU cloud providers. High-performance computing and distributed systems. Training optimisation, memory management and networking. MLOps tooling ...

Senior Engineer, Systematic Data

Hiring Organisation
Balyasny Asset Management
Location
London, United Kingdom
Salary
£ 100 K
programming languages (e.g. Java/C#) Experience with datasets used for cash equity research and trading Experience with distributed computing (e.g. Spark/Dask, Slurm/HTCondor) Experience with event-driven, asynchronous architectures and related messaging technologies (e.g. Kafka, MQ) Experience with cloud platforms (e.g. AWS, GCP, Azure) Experience ...

Senior HPC Systems Administrator

Hiring Organisation
University of Oxford
Location
Oxford, Oxfordshire, United Kingdom
Salary
£ 50 K
troubleshooting and systems integration skills● Excellent written and verbal communication abilities● Ability to work independently and collaboratively in a research environmentDesirable Skills● Experience with SLURM or other job scheduling systems● Experience with containerisation and virtualisation technologies● Knowledge of cloud computing platforms● Familiarity with scientific computing tools and AI/ ...

HPC Systems Engineer

Location
Greater London, England, United Kingdom
Need: At least 5+ years of professional experience in high performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required At least 5+ years of experience with Linux systems administration High proficiency ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, United Kingdom
Salary
£ 80 K
scale distributed training or HPC systems.Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.Experience working with containerized environments (Docker, Singularity/ ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, UK
Employment Type
Full-time
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity ...

Hybrid HPC System Administrator – Slurm, Linux & Storage

Location
Farnborough, England, United Kingdom
Lenovo in the United Kingdom (Hampshire) is hiring an HPC System Administrator to manage data center infrastructure, Linux/Unix environments, and monitoring to ensure uptime and security. This role blends on-site and remote ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Senior Data & MLOps Engineer

Location
Greater London, England, United Kingdom
least one systems language (Go, Rust, or C++). Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm). Familiarity with GPU telemetry systems such as NVML or DCGM and hardware‐level monitoring concepts. Demonstrated experience scaling systems from Proof‐of‐Concept ...

Senior IT Systems Administrator (Linux & HPC)

Hiring Organisation
KBR
Location
Guildford, Surrey, United Kingdom
Salary
£ 80 K
reliability of Linux-based enterprise and high-performance computing platforms. The role will provide hands-on technical ownership across Linux operating systems, HPC compute, workload scheduling and associated storage and network services.A major focus will be the operation of cloud and co-located HPC platforms. The position requires strong … Linux engineering experience, practical knowledge of HPC environments, including NVIDIA Base Command Manager, Azure CycleCloud and SLURM. The ability to diagnose complex cross-platform issues, and the confidence to work independently while acting as a senior technical resource for colleagues is also essential. This role does not include formal ...

Senior Field Engineer - AI/ML HPC & Kubernetes Solutions

Location
Greater London, England, United Kingdom
drive proofs of concept to accelerate client deployments. You will collaborate with engineering teams, contribute to product direction, and help optimize workloads using Slurm, NCCL, and Infiniband in high-performance environments. #J-18808-Ljbffr ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
agency engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation.Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential.Relevant Background: Experience ...