1 to 25 of 33 Slurm Workload Manager Jobs in the UK

Member of Technical Staff (AI Infrastructure Engineer)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Full time Location Type Hybrid Department AI We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams … scale AI training and inference clusters. Responsibilities Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement ...

Senior System Engineer (Munich, Germany)

Hiring Organisation
Jobleads-UK
Location
Cambourne, England, United Kingdom
debugging complex distributed systems where the root cause could be code, network, or silicon. Preferred qualifications HPC Background: Experience working with traditional supercomputing schedulers (Slurm, PBS) or modern batch schedulers (Volcano, Kueue, Ray). Bare Metal Provisioning: Experience with tools like Cluster API (CAPI), Metal3, Tinkerbell, Canonical MaaS ...

Machine Learning Systems & Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
serving: Operate the systems researchers use to launch experiments, data jobs, and production endpoints — workflow engines (e.g., Kubeflow Pipelines, Airflow), GPU schedulers (e.g., Volcano, Slurm), experiment trackers (e.g., MLflow, Weights & Biases), and managed‐inference platforms (e.g., Modal, Triton) — and maintain a launcher SDK for one‐command runs. Containerization ...

Infrastructure / DevOps Lead

Hiring Organisation
Jobleads-UK
Location
United Kingdom
details. Nice to Have Experience managing physical data centres, co‐location facilities, or hybrid infrastructure environments. Working knowledge of ML orchestration frameworks (e.g., Ray, Slurm, Kubeflow). Background in media pipelines, VFX tooling, or media compliance standards (MPA, ISO 27001). Prior experience working in a hybrid startup/ ...

Enterprise Architect - AI

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Founding AI Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technical direction of the company from day one. What we’re looking for Large‐scale distributed training. PyTorch and modern deep‐learning frameworks. Kubernetes, Slurm or GPU orchestration platforms. AWS and specialist GPU cloud providers. High‐performance computing and distributed systems. Training optimisation, memory management and networking. MLOps tooling ...

ML Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Claude Code, Codex, Kimi Code, Pi Agent, Droid, or similar agentic coding systems as a development surface Experience with GPU clusters on Kubernetes, Slurm, Ray, custom schedulers, or cloud GPU orchestration NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, or EFA Rust, C++, CUDA, Go, or systems‐level performance work ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation. Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands‐on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential. Relevant Background: Experience ...

Research Engineer, Pre-Training

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Background in numerical computing, HPC, or distributed systems, including familiarity with GPUs/TPUs, high-performance networking (NVLink/InfiniBand), Kubernetes/Slurm, and OS internals Expertise in Python and deep experience with modern deep learning frameworks (PyTorch and/or JAX) Advanced degree (MS or PhD) in Computer ...

Senior Cloud Platform Engineer—AI Infra & Automation

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU). Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Strong skillset and experience in end‐to‐end deployment automation and CI of containerised services. Complete automation … pipelines for build, test, deploy, manage, alert, destroy, rebuild. Experience with managing production Kubernetes clusters and workloads. Experience with workload queue management systems (SLURM, LSF, Kueue). Experience with managed switch configuration (e.g. EOS, SONiC, DNOS). Programming experience with Python3 utilising classes and inheritance. Programming experience with ...

Principal Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU). Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Strong skillset and experience in end‐to‐end deployment automation and CI of containerised services. Complete automation … pipelines for build, test, deploy, manage, alert, destroy, rebuild. Experience with managing production Kubernetes clusters and workloads. Experience with workload queue management systems (SLURM, LSF, Kueue). Experience with managed switch configuration (e.g. EOS, SONiC, DNOS). Programming experience with Python3 utilising classes and inheritance. Programming experience with ...

Senior AI Infrastructure Engineer - Scale Multi-GPU Training

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
will architect and optimize distributed training across multiple GPUs and machines in AWS, eliminate bottlenecks in the data path, and manage cluster orchestration with Slurm and Kubernetes. The role requires deep PyTorch expertise, familiarity with transformer models, and experience deploying production AI systems. #J-18808-Ljbffr ...

AI infrastructure engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
will eliminate bottlenecks in the data path to ensure training is fast and as capital efficient as possible alongside managing cluster orchestration using slurm and Kubernetes while preparing to expand into specialised GPU providers. And finally you will master the stack from pytorch based learning libraries to complex data ...

Research Engineer, Machine Learning – Paris/London/Zurich/Warsaw

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
+ years working on large‐scale ML codebases. Hands‐on with PyTorch, JAX or TensorFlow; comfortable with distributed training (DeepSpeed/FSDP/SLURM/K8s). Experience in deep learning, NLP or LLMs; bonus for CUDA or data‐pipeline chops. Strong software‐design instincts: testing, code review ...

Research Engineer/Scientist - Machine Learning RL & Optimisation (Contractor)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Monte Carlo Tree Search (MCTS). Deep understanding of GPU architectures, kernels, FlashAttention, and profiling tools. Familiarity with cluster environments and schedulers like Slurm or Kubernetes. Hands-on experience developing within NPU stack or ecosystem integrations. Practical knowledge of hardware-expressive Domain-Specific Languages (e.g., TileLang, Triton) to optimize ...

RF Signature Analyst

Hiring Organisation
MASS Consultants
Location
Fareham, Hampshire, South East, United Kingdom
Employment Type
Permanent
Salary
£55,000
SolidWorks or RhinoCAD STEM degree It would be great if you also have: Understanding of Linux environments and High-Performance Computing (HPC) systems, including Slurm Understanding of advanced combat air, weapons or UAS (Uncrewed Air Systems) capabilities Experience within the defence sector and/or an understanding of survivability ...

AI Inference Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

Quant Developer (C++)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
global team Python, Q/kdb+ Testing methodologies (unit tests, regression tests) Dev workflow – SVN, GIT, JIRA, Code Reviews, etc Grid & cluster tools (especially SLURM) The minimum base salary for this role is $120,000 if located in New York. This expectation is based on available information ...

Senior Cloud Networking Engineer — High‐Performance Infra

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
focus on end‐user availability. Desirable but not required Experience with Openstack cloud platform s. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Experience with hardware offloading on RDMA‐capable NICs and how that integrates with virtual networking on Open V‐switch ...

Software Engineer, General

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
observability tools (Prometheus, Grafana) is beneficial Bonus/Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems‐Level ...

High Performance Computer Scientist /HPC Developer

Hiring Organisation
IT Graduate Recruitment
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
£50,000 per annum
large-scale distributed systems. Research experience involving computational workloads. Experience with parallel programming (MPI, OpenMP, CUDA). Knowledge of scheduling systems such as Slurm, PBS or LSF. Contributions to technical projects, open source or research communities. Experience working with advanced computing environments. Academic Focus We are particularly interested … Parallel Programming, MPI, OpenMP, Multithreading, Concurrency, Algorithms, Data Structures, Systems Design, Kernel Development, Networking, Storage Systems, Distributed Storage, Automation, Infrastructure Automation, Shell Scripting, Bash, Slurm, PBS, LSF, Workload Scheduling, Resource Management, Linux Administration, Server Infrastructure, Cloud Infrastructure, AWS HPC, Azure HPC, Data Processing, Machine Learning Infrastructure, AI Infrastructure ...

MLOps Engineer

Hiring Organisation
Jobleads-UK
Location
Oxford, England, United Kingdom
across on-premises accelerator clusters and cloud (GPU/CPU) for training and simulation workloads Drive infrastructure-as-code practices: containerisation, orchestration (Kubernetes/Slurm), and reproducible environment management Contribute to the internal developer platform: self-service tooling, documentation, and runbooks that raise engineering productivity across the company What … Experience with experiment tracking and model lifecycle management tools (MLflow, W&B, DVC, or similar) Solid understanding of containerisation (Docker) and orchestration (Kubernetes or Slurm) for distributed compute workloads Infrastructure-as-code mindset: Terraform, Ansible, or equivalent; CI/CD pipelines (GitHub Actions, Jenkins, or similar) Experience with hardware ...

Solution Architect - GPU & HPC

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
winning them requires more than a great sales team. The Solutions Architect sits at the intersection of sales, infrastructure, and the customer, translating complex workload requirements into technically sound, commercially viable solutions on the Hyperstack platform. You’ll be the primary technical authority through the sales cycle: engaging directly … proposal, and delivery handover — acting as the primary technical authority for GPU cloud solution design. Engage directly with prospective and existing customers to understand workload requirements, technical constraints, and commercial objectives, producing detailed solution designs including architecture diagrams, network topology, storage configurations, and GPU resource allocation models. Collaborate closely ...

Senior Cloud Engineer (K8S)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
cloud platforms. Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary, Graphcore offers flexible ...

Senior Software Engineer, Inference Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
model lifecycle management is highly desirable Bonus/Good to Have HPC & Cluster Management: Experience handling large‐scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large‐scale data processing frameworks Systems-Level ...