1 to 25 of 26 Slurm Workload Manager Jobs in the UK

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 100 K
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub.Knowledge of other programming ...

Scientific Computing Engineer - Vernalis

Hiring Organisation
Vernalis
Location
Cambridgeshire, England, United Kingdom
configurations for the Enterprise IT team. Key Responsibilities HPC Management: Oversee the health and performance of our on-premises HPC cluster ( Bright Computing/SLURM ). Optimise job scheduling and resource allocation across mixed CPU and GPU nodes for computational chemistry and ML applications. DevOps & Container Orchestration: Lead … Linux (RHEL) and SELinux security policies. DevOps: Strong experience with GitLab CI/CD and container orchestration (Docker Swarm). HPC Stack: Proficiency in SLURM workload management and Bright Computing cluster tools. Storage & Identity: Experience with NFS, enterprise storage (EMC2), and managing UID/GID/SID mappings ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 70 K
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads.AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms.ML frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent … ArgoCD) for continuous, declarative platform delivery.Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines.Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

MLOps Engineer

Hiring Organisation
Intellectual Capital Resources
Location
Oxford, Oxfordshire, United Kingdom
Salary
£ 80 K
pipelines ML infrastructure expertise- experiment tracking, model management, and deployment Strong Python skills- experience with PyTorch, JAX, or similar frameworks Platform engineering- Docker, Kubernetes, Slurm, and distributed systems CI/CD knowledge- automating testing, validation, and deployments Performance optimisation- profiling and benchmarking AI workloads Infrastructure as Code- Terraform, Ansible ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
titles matter less than experience. We're particularly interested in people with experience in:Large-scale distributed training.PyTorch and modern deep learning frameworks.Kubernetes, Slurm or GPU orchestration platforms.AWS and specialist GPU cloud providers.High-performance computing and distributed systems.Training optimisation, memory management and networking.MLOps tooling, CI/CD and experimentation ...

HPC Specialist Architect - Energy Industry

Hiring Organisation
AmazonWebServices
Location
London, United Kingdom
Salary
£ 80 K
more of the following programming languages: C++, Python, Cuda, or Bash.- Experience in architecting an HPC platform with scheduling middleware (e.g. Slurm, Torque, Symphony or GridServer) and in deployment, tuning and management of HPC technologies in a multi-user environment.- High level understanding of the underlying infrastructure platform ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
agency engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation.Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential.Relevant Background: Experience ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
reliability, and cost-efficiency.Your ImpactProvision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible.Optimise GPU compute workloads, high-speed networking … from individuals. Anaplan does not: Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person. Send job offers via email. All offers are first extended verbally by a member of our internal recruitment team ...

HPC Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 60 K
This company is on the hunt for HPC Engineers to power their 25 Petabyte system.... Sound good? Well there's more! Imagine working with Slurm clusters and GPFS storage, all while being an integral part of groundbreaking translational research. You will work in a dynamic team of five, where ...

AI infrastructure engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 100 K
throughput, and ensuring training is as efficient and cost-effective as possible. You'll also play a critical role in managing cluster orchestration with Slurm and Kubernetes while helping evolve the platform to support next-generation GPU infrastructure and specialised compute providers. This is an opportunity to work across ...

AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 80 K
will eliminate bottlenecks in the data path to ensure training is fast and as capital efficient as possible alongside managing cluster orchestration using slurm and Kubernetes while preparing to expand into specialised GPU providers. And finally you will master the stack from pytorch based learning libraries to complex data ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
workloads.Exposure to multi-tenant serving or SLA-driven infrastructure.Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.Familiarity with Kubernetes/Slurm for cluster orchestration.Interest or experience in energy markets, grid systems, or sustainability-focused compute.BenefitsCompetitive salary and an equity sign-on bonus.Biannual bonus scheme.Fully expensed ...

Senior Applied Research Engineer - Video Team

Hiring Organisation
Synthesia
Location
London, United Kingdom
Salary
£ 80 K
human-centric generationFamiliarity with world/interactive modelsExperience with GANs or VAEsExperience optimizing inference systems for productionOur stackPython, PyTorch, CUDADeepSpeed, distributed training & inferenceSequence parallelismAWS, SLURM, DockerGitHub, CI/CD pipelinesWho you areYou are research-driven but outcome-focusedYou care about shipping, not just publishingYou can explore multiple ideas quickly ...

Principal Cloud Architect – HPC/GPU & AI Platform Solutions

Hiring Organisation
Oracle Corporation
Location
London, United Kingdom
Salary
£ 80 K
public cloud platforms. Infrastructure as Code using Terraform and Ansible. Python, Bash, and PowerShell scripting. Kubernetes and container orchestration. HPC cluster management platforms including Slurm, PBS, or Bright Cluster Manager. High-performance networking technologies including RDMA and InfiniBand. MPI and distributed file systems. Cloud-native architectures. AI/… Large Language Models (LLMs) Agentic AI AI Platform Architecture Inference Serving Programming & AutomationPython Bash PowerShell Automation Frameworks Infrastructure Automation HPC TechnologiesSlurm PBS Bright Cluster Manager RDMA InfiniBand MPI Distributed File Systems Customer & ConsultingSolution Architecture Technical Consulting Pre-Sales Executive Presentations Customer Workshops Technical Enablement AI Transformation Strategy Cloud Adoption ...

HPC Senior Technology Consultant

Hiring Organisation
Hewlett Packard Enterprise
Location
Wokingham, Berkshire, United Kingdom
Salary
£ 70 K
customer-facing environment.Desirable ExperienceExperience in one or more of the following areas would be advantageous:Cluster management platforms such as Bright Cluster Manager, HPE Performance Cluster Manager (HPCM) or Cray Systems Manager (CSM).SLURM or PBS Pro job schedulers.High-performance networking technologies such as InfiniBand or Slingshot.Parallel ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, United Kingdom
Salary
£ 80 K
scalable engineering solutions.Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows.Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability, and reliability … NVIDIA Nsight, DCGM, and related ecosystem technologies.Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation.Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks.Demonstrated success bridging research and production environments, enabling rapid experimentation while maintaining operational ...

Build Engineer

Hiring Organisation
Intellectual Capital Resources
Location
London, United Kingdom
Salary
£ 80 K
hybrid working model, required onsite 3 days a week. Experience for the Build Engineer includes: Software development with Python Modern build systems, e.g. Bazel Workload management, e.g. Slurm, LSF or SGE Infrastructure as code or IAC, e.g. Ansible or Terraform Containerisation, e.g. Docker Desired: Experience supporting silicon ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
NTT
Location
London, United Kingdom
Salary
£ 80 K
land service-led AI solutions. Lead integration of compute, storage, networking, the AI software stack (CUDA, ROCm, Triton, NIM, NVIDIA AI Enterprise, Run:ai, Slurm, Kubernetes/Kubeflow) and managed-service operating models across multiple domains, delivery units and geographies. Build business cases, TCO and unit-economics models (cost … tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power, liquid cooling ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
The Nippon Telegraph And Telephone Corporation (NTT)
Location
United Kingdom
Salary
£ 70 K
land service-led AI solutions. Lead integration of compute, storage, networking, the AI software stack (CUDA, ROCm, Triton, NIM, NVIDIA AI Enterprise, Run:ai, Slurm, Kubernetes/Kubeflow) and managed-service operating models across multiple domains, delivery units and geographies. Build business cases, TCO and unit-economics models (cost … tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power, liquid cooling ...

Python Software Engineer - Intraday Trading

Hiring Organisation
Millennium Management
Location
London, United Kingdom
Salary
£ 100 K
strong understanding of Linux operating systems• Experience with grid scheduling and compute orchestration for real-time compute management at scale, including technologies such as SLURM• Strong understanding of event-driven architecture and experience with messaging and caching technologies such as Kafka, Solace, Pulsar, Memcache, and Redis• Experience building … scale real-time portfolio analytics tools, cloud platforms, containerization technologies such as Docker and Kubernetes, or multi-threaded C++ is a plusRecruiter:Ruby KazmiHiring Manager:Shashank GiriDepartment:Information Technology ...

Staff Software Engineer, Kubernetes Platform

Hiring Organisation
Humanloop
Location
London, United Kingdom
Salary
> £ 150 K
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure.We make sure the control plane is fast, correct, and always available. … Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemptionScale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds usDesign, build, and operate core cluster services such ...

Founding GPU Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
center systems. You'll work on low-level performance engineering for large-scale compute clusters, helping Fuse build the software layer that ties GPU workload behaviour to energy availability and grid demand.The OpportunityDemand for high-performance compute capacity across the markets we operate in significantly outpaces what … multi-node scaling using NCCL, MPI, or similar communication libraries.Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity.Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines.Benchmark against ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Part Time, Work From Home
Salary
From £1,000 to £1,200 per day
NVSwitch, InfiniBand and RoCE. Designing high-performance networking and storage environments. Supporting both bare-metal and Kubernetes-based deployments. Working with technologies such as Slurm, Kubernetes, NVIDIA GPU Operator and NCCL. Supporting technical sales opportunities and major customer engagements. Running technical workshops with senior customer stakeholders. Supporting RFP/… Experience with Blackwell, HGX, DGX, NVLink/NVSwitch, InfiniBand or RoCE would be highly advantageous. Were also interested in people with experience across: Kubernetes Slurm NVIDIA GPU Operator NCCL RDMA/GPUDirect High-performance storage GPU cluster networking HLD/LLD production Technical pre-sales and customer-facing architecture ...

HPC Operations Lead

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 70 K
operation of high performance storage services, supporting both internal workloads and external collaboration. The environment includes large scale HPC clusters, Linux based systems, workload schedulers such as Slurm, networking with Infiniband and parallel file systems such as GPFS. Experience with high performance storage at petabyte scale is particularly ...