1 to 25 of 85 Slurm Workload Manager Jobs in the UK

HPC Engineer - Slurm Expertise

Hiring Organisation
TEKsystems
Location
Stevenage, Hertfordshire, UK
Employment Type
Full-time
Description Job Title: HPC Engineer Job Description We are seeking an experienced Senior HPC Engineer with a core expertise in the Slurm workload manager. This role is pivotal in owning scheduling, job and queue design, and cluster policy management across our HPC estate. The primary mandate of this … position is Slurm ownership. Responsibilities Own and manage the Slurm workload manager. Design and implement job and queue scheduling strategies. Develop and enforce cluster policies across the HPC estate. Manage, secure, and maintain Red Hat Enterprise Linux (RHEL) infrastructure across versions 7, 8, and 9. Provide support ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Machine Learning Engineer (Mid to Principal)

Location
West of England, England, United Kingdom
Jira. Edge deployments: Nvidia Jetson (e.g. AGX Orin), Raspberry Pi, or other embedded accelerators. Distributed model training & infra: Pytorch DDP, FDSP and TorchTitan, Megatron, Slurm, Run:ai, DeepSpeed, Kubernetes, cloud or on-prem GPU clusters. About you You’ve built ML systems that persist—deployed in real settings, iterated ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub. Knowledge of other ...

Infrastructure Support Engineer - Contract

Location
Greater London, England, United Kingdom
working in a ticket queue against response and resolution targets. Clear written runbooks and documentation. Useful : Cloudflare Zero Trust and WARP. FlexLM licence administration. SLURM and HPC or EDA compute environments. Semiconductor or engineering environments using Cadence or Synopsys tools. Kubernetes. Due to U.S. export control regulations, candidates' eligibility ...

HPC System Administrator

Location
Farnborough, England, United Kingdom
defects. Perform Software and firmware testing for any fixes, upgrades, security patch. Position requirements: Hands-on experience with HPC tools and schedulers such as Slurm and PBS Pro, as well as Confluent Solid scripting and automation skills using Ansible, Bash, and Python Experience with distributed storage systems and file ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
environments Enjoys solving complex technical problems collaboratively Nice-to-Haves Experience with large scale model training or distributed inference systems Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms Experience with InfiniBand, RDMA, or high-performance networking Experience operating bare metal infrastructure Familiarity with storage systems commonly used ...

Senior HPC Systems Administrator

Location
Oxford, England, United Kingdom
systems integration skills Excellent written and verbal communication abilities Ability to work independently and collaboratively in a research environment Desirable Skills Experience with SLURM or other job scheduling systems Experience with containerisation and virtualisation technologies Knowledge of cloud computing platforms Familiarity with scientific computing tools and AI/ ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
matter less than experience. We're particularly interested in people with experience in: Large-scale distributed training. PyTorch and modern deep learning frameworks. Kubernetes, Slurm or GPU orchestration platforms. AWS and specialist GPU cloud providers. High-performance computing and distributed systems. Training optimisation, memory management and networking. MLOps tooling ...

Senior Engineer, Systematic Data

Hiring Organisation
Balyasny Asset Management
Location
London, UK
Employment Type
Full-time
programming languages (e.g. Java/C#) Experience with datasets used for cash equity research and trading Experience with distributed computing (e.g. Spark/Dask, Slurm/HTCondor) Experience with event-driven, asynchronous architectures and related messaging technologies (e.g. Kafka, MQ) Experience with cloud platforms (e.g. AWS, GCP, Azure) Experience ...

Bioinformatician - Pathogen Pathogen Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
clinical metagenomics or targeted sequencing. Exposure to cloud computing platforms (e.g. OCI, AWS or GCP) and/or high‐performance computing (HPC) environments (e.g. Slurm) for running and managing bioinformatics analyses. Experience using version control (e.g. Git) to develop, maintain and share reproducible bioinformatics code and workflows within ...

Data Engineer (Forward Deployed)

Location
Oxford, England, United Kingdom
environments using containers (e.g. Docker). Scale processing across distributed cloud warehouses/storage via container orchestration (e.g. Kubernetes), High-Performance GPU Compute (e.g. Slurm), or distributed compute frameworks (e.g. Spark, Ray). Contribute to an engineering culture that values maintainability, testing, robust system design, and deep collaboration ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, UK
Employment Type
Full-time
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity ...

Lead Infrastructure Engineer - Linux

Hiring Organisation
Sopra Steria
Location
United Kingdom
Employment Type
Permanent
Salary
GBP Annual
large-scale Linux environments, ensuring high availability, performance, security, and resilience. Architecting and supporting High Performance Computing (HPC) platforms, including compute, storage, networking, workload scheduling, and user environments. Designing, deploying, and managing container and orchestration platforms to support modern application delivery. Acting as a Linux Infrastructure Solutions Architect, driving … platform technologies. What you'll bring: Extensive experience administering enterprise Linux environments (RHEL, Rocky Linux, Ubuntu, or equivalent). Strong knowledge of HPC infrastructure, workload management ideally using Slurm, and performance optimisation. Proven experience designing and delivering Linux infrastructure and platform architectures. Strong understanding of Linux security, virtualisation ...

Hybrid HPC System Administrator – Slurm, Linux & Storage

Location
Farnborough, England, United Kingdom
Lenovo in the United Kingdom (Hampshire) is hiring an HPC System Administrator to manage data center infrastructure, Linux/Unix environments, and monitoring to ensure uptime and security. This role blends on-site and remote ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation. Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential. Relevant Background: Experience ...

Senior Research Software Engineer - Breakthrough Listen

Location
Oxford, England, United Kingdom
shell scripting; implementation, deployment and management of large (multi-billion row) relational databases; containers and orchestration; AI-assisted software development; distributed systems (e.g. Slurm, Kubernetes) and high performance computing. This post is offered as an open-ended position subject to external funding. It is full time, but flexible options ...

Research Engineer, Forge

Location
Greater London, England, United Kingdom
comfort in fast‐moving, under‐specified environments. Nice to have Distributed training experience (FSDP, DeepSpeed, Megatron, etc.). Cluster/orchestration experience (SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, etc.). Experience building reliable ML infrastructure, evaluation systems, or large‐scale data processing pipelines. Research experience in LLMs, agents, multimodal ...

Senior Data & MLOps Engineer

Location
Greater London, England, United Kingdom
least one systems language (Go, Rust, or C++). Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm). Familiarity with GPU telemetry systems such as NVML or DCGM and hardware‐level monitoring concepts. Demonstrated experience scaling systems from Proof‐of‐Concept ...

Cloud Engineer

Location
West of England, England, United Kingdom
openstack HPC stackhpc community news AI azimuth ironic monitoring deployment baremetal kolla-ansible ansible networking kayobe kolla infrastructure kubernetes slurm monasca scheduling newsletter mpi virtualisation beegfs cloudkitty sriov cluster training ciuk operations recruitment network overcloud slinky ceph data github ucx galaxy cloud-init storage gpu grafana cuda vm kata … following additional technologies is always beneficial: O/S : Rocky Linux, Ubuntu Storage : Ceph, NFS Networking : Ethernet, Infiniband, SRIOV Compute and Workloads : Kubernetes, Slurm Our technology stack is constantly evolving to meet market and user requirements. So we will help you get up to speed as required. Flexible working ...

Senior IT Systems Administrator (Linux & HPC)

Location
Leatherhead, England, United Kingdom
reliability of Linux-based enterprise and high-performance computing platforms. The role will provide hands-on technical ownership across Linux operating systems, HPC compute, workload scheduling and associated storage and network services.A major focus will be the operation of cloud and co-located HPC platforms. The position requires strong … Linux engineering experience, practical knowledge of HPC environments, including NVIDIA Base Command Manager, Azure CycleCloud and SLURM. The ability to diagnose complex cross-platform issues, and the confidence to work independently while acting as a senior technical resource for colleagues is also essential. This role does not include formal ...

Lead Linux Engineer

Hiring Organisation
Sopra Steria
Location
Salisbury, Wiltshire, South West, United Kingdom
Employment Type
Permanent
Salary
25 days holidays, 6% contributory pension, 4 x life insurance, 3% Flex
large-scale Linux environments, ensuring high availability, performance, security, and resilience. Architecting and supporting High Performance Computing (HPC) platforms, including compute, storage, networking, workload scheduling, and user environments. Designing, deploying, and managing container and orchestration platforms to support modern application delivery. Acting as a Linux Infrastructure Solutions Architect, driving … platform technologies. What you'll bring: Extensive experience administering enterprise Linux environments (RHEL, Rocky Linux, Ubuntu, or equivalent). Strong knowledge of HPC infrastructure, workload management ideally using Slurm, and performance optimisation. Proven experience designing and delivering Linux infrastructure and platform architectures. Strong understanding of Linux security, virtualisation ...

HPC Systems Engineer: Research Compute & AI Workloads

Location
Greater London, England, United Kingdom
support a large scale computing environment used for data-intensive research, modelling and AI workloads. You will work across Linux compute environments, GPU infrastructure, Slurm, and containerisation with Docker or Singularity, driving automation and platform improvements to better serve researchers. #J-18808-Ljbffr ...

Senior Field Engineer - AI/ML HPC & Kubernetes Solutions

Location
Greater London, England, United Kingdom
drive proofs of concept to accelerate client deployments. You will collaborate with engineering teams, contribute to product direction, and help optimize workloads using Slurm, NCCL, and Infiniband in high-performance environments. #J-18808-Ljbffr ...