1 to 25 of 56 Slurm Workload Manager Jobs in the UK

Member of Technical Staff (AI Infrastructure Engineer)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Full time Location Type Hybrid Department AI We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams … scale AI training and inference clusters. Responsibilities Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement ...

Senior System Engineer (Munich, Germany)

Hiring Organisation
Jobleads-UK
Location
Cambourne, England, United Kingdom
debugging complex distributed systems where the root cause could be code, network, or silicon. Preferred qualifications HPC Background: Experience working with traditional supercomputing schedulers (Slurm, PBS) or modern batch schedulers (Volcano, Kueue, Ray). Bare Metal Provisioning: Experience with tools like Cluster API (CAPI), Metal3, Tinkerbell, Canonical MaaS ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 100 K
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub.Knowledge of other programming ...

Machine Learning Systems & Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
serving: Operate the systems researchers use to launch experiments, data jobs, and production endpoints — workflow engines (e.g., Kubeflow Pipelines, Airflow), GPU schedulers (e.g., Volcano, Slurm), experiment trackers (e.g., MLflow, Weights & Biases), and managed‐inference platforms (e.g., Modal, Triton) — and maintain a launcher SDK for one‐command runs. Containerization ...

Infrastructure / DevOps Lead

Hiring Organisation
Jobleads-UK
Location
United Kingdom
details. Nice to Have Experience managing physical data centres, co‐location facilities, or hybrid infrastructure environments. Working knowledge of ML orchestration frameworks (e.g., Ray, Slurm, Kubeflow). Background in media pipelines, VFX tooling, or media compliance standards (MPA, ISO 27001). Prior experience working in a hybrid startup/ ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 70 K
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

Cluster Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
with quality, reliability, safety, and performance goals in mind. Work with the Software Engineering and Solutions Engineering teams to ensure the infrastructure supports AI workload requirements. Engage with project management to define milestones, risks, resourcing, and deliverables associated with the delivery. Produce accurate architecture diagrams, configuration documentation, runbooks ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads.AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms.ML frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent … ArgoCD) for continuous, declarative platform delivery.Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines.Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

MLOps Engineer

Hiring Organisation
Intellectual Capital Resources
Location
Oxford, Oxfordshire, United Kingdom
Salary
£ 80 K
pipelines ML infrastructure expertise- experiment tracking, model management, and deployment Strong Python skills- experience with PyTorch, JAX, or similar frameworks Platform engineering- Docker, Kubernetes, Slurm, and distributed systems CI/CD knowledge- automating testing, validation, and deployments Performance optimisation- profiling and benchmarking AI workloads Infrastructure as Code- Terraform, Ansible ...

Senior Infrastructure Engineer, Research Singapore

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Enterprise Architect - AI

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Founding AI Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technical direction of the company from day one. What we’re looking for Large‐scale distributed training. PyTorch and modern deep‐learning frameworks. Kubernetes, Slurm or GPU orchestration platforms. AWS and specialist GPU cloud providers. High‐performance computing and distributed systems. Training optimisation, memory management and networking. MLOps tooling ...

Staff ML Infrastructure Engineer August 4, 2026

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
. Experience with cloud infrastructure such as AWS & GCP . Experience operating distributed systems in production. Beneficial Skills Argo Workflows or Kubernetes workflow engines. SLURM or other HPC job schedulers. ML experiment tracking tools such as Weights & Biases or MLflow. Data versioning or lakehouse technologies such as LakeFS, Iceberg ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
titles matter less than experience. We're particularly interested in people with experience in:Large-scale distributed training.PyTorch and modern deep learning frameworks.Kubernetes, Slurm or GPU orchestration platforms.AWS and specialist GPU cloud providers.High-performance computing and distributed systems.Training optimisation, memory management and networking.MLOps tooling, CI/CD and experimentation ...

ML Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Claude Code, Codex, Kimi Code, Pi Agent, Droid, or similar agentic coding systems as a development surface Experience with GPU clusters on Kubernetes, Slurm, Ray, custom schedulers, or cloud GPU orchestration NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, or EFA Rust, C++, CUDA, Go, or systems‐level performance work ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker ...

HPC Operations Engineer - Banking & Finance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

HPC Specialist Architect - Energy Industry (AWS)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
more of the following programming languages: C++, Python, Cuda, or Bash. Experience in architecting an HPC platform with scheduling middleware (e.g., Slurm, Torque, Symphony or GridServer) and in deployment, tuning and management of HPC technologies in a multi‐user environment. High level understanding of the underlying infrastructure platform ...

Research Engineer, Pre-Training

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Background in numerical computing, HPC, or distributed systems, including familiarity with GPUs/TPUs, high-performance networking (NVLink/InfiniBand), Kubernetes/Slurm, and OS internals Expertise in Python and deep experience with modern deep learning frameworks (PyTorch and/or JAX) Advanced degree (MS or PhD) in Computer ...

HPC Specialist Architect - Energy Industry

Hiring Organisation
AmazonWebServices
Location
London, United Kingdom
Salary
£ 80 K
more of the following programming languages: C++, Python, Cuda, or Bash.- Experience in architecting an HPC platform with scheduling middleware (e.g. Slurm, Torque, Symphony or GridServer) and in deployment, tuning and management of HPC technologies in a multi-user environment.- High level understanding of the underlying infrastructure platform ...

Energy HPC Architect: AWS Solutions & PoC Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
more of the following programming languages: C++, Python, Cuda, or Bash. - Experience in architecting an HPC platform with scheduling middleware (e.g. Slurm, Torque, Symphony or GridServer) and in deployment, tuning and management of HPC technologies in a multi-user environment. - High level understanding of the underlying infrastructure platform ...

Research Engineer/Scientist - Machine Learning RL & Optimisation (Contractor)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Monte Carlo Tree Search (MCTS). Deep understanding of GPU architectures, kernels, FlashAttention, and profiling tools. Familiarity with cluster environments and schedulers like Slurm or Kubernetes. Hands-on experience developing within NPU stack or ecosystem integrations. Practical knowledge of hardware-expressive Domain-Specific Languages (e.g., TileLang, Triton) to optimize ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
agency engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation.Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential.Relevant Background: Experience ...

Research Engineer, Machine Learning – Paris/London/Zurich/Warsaw

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
+ years working on large‐scale ML codebases. Hands‐on with PyTorch, JAX or TensorFlow; comfortable with distributed training (DeepSpeed/FSDP/SLURM/K8s). Experience in deep learning, NLP or LLMs; bonus for CUDA or data‐pipeline chops. Strong software‐design instincts: testing, code review ...

RF Signature Analyst

Hiring Organisation
MASS Consultants
Location
Fareham, Hampshire, South East, United Kingdom
Employment Type
Permanent
Salary
£55,000
SolidWorks or RhinoCAD STEM degree It would be great if you also have: Understanding of Linux environments and High-Performance Computing (HPC) systems, including Slurm Understanding of advanced combat air, weapons or UAS (Uncrewed Air Systems) capabilities Experience within the defence sector and/or an understanding of survivability ...