1 to 25 of 37 Slurm Workload Manager Jobs in London

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 100 K
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub.Knowledge of other programming ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub. Knowledge of other ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 70 K
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads.AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms.ML frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent … ArgoCD) for continuous, declarative platform delivery.Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines.Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms. ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron … continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Enterprise Architect - AI

Location
Greater London, England, United Kingdom
continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
titles matter less than experience. We're particularly interested in people with experience in:Large-scale distributed training.PyTorch and modern deep learning frameworks.Kubernetes, Slurm or GPU orchestration platforms.AWS and specialist GPU cloud providers.High-performance computing and distributed systems.Training optimisation, memory management and networking.MLOps tooling, CI/CD and experimentation ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
matter less than experience. We're particularly interested in people with experience in: Large-scale distributed training. PyTorch and modern deep learning frameworks. Kubernetes, Slurm or GPU orchestration platforms. AWS and specialist GPU cloud providers. High-performance computing and distributed systems. Training optimisation, memory management and networking. MLOps tooling ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Senior kdb+ Developer, Vice President

Hiring Organisation
State Street Bank
Location
London, United Kingdom
Salary
£ 80 K
Python.Proficiency with modern data science toolchains, including Jupyter, pandas, NumPy, scikit‐learn, with applied machine‐learning experience.Solid understanding of parallel computing frameworks such as Slurm or equivalent technologies.Strong background in quantitative analysis, including mathematical modelling, statistics, regression, and probability theory.Bachelor’s or Master’s degree in Computer Science, Mathematics ...

Senior kdb+ Developer, Vice President

Hiring Organisation
State Street Bank
Location
London, UK
Employment Type
Full-time
with modern data science toolchains, including Jupyter, pandas, NumPy, scikit‐learn, with applied machine‐learning experience. Solid understanding of parallel computing frameworks such as Slurm or equivalent technologies. Strong background in quantitative analysis, including mathematical modelling, statistics, regression, and probability theory. Bachelor's or Master's degree in Computer ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
reliability, and cost-efficiency.Your ImpactProvision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible.Optimise GPU compute workloads, high-speed networking … from individuals. Anaplan does not: Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person. Send job offers via email. All offers are first extended verbally by a member of our internal recruitment team ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
agency engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation.Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential.Relevant Background: Experience ...

HPC Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 60 K
This company is on the hunt for HPC Engineers to power their 25 Petabyte system.... Sound good? Well there's more! Imagine working with Slurm clusters and GPFS storage, all while being an integral part of groundbreaking translational research. You will work in a dynamic team of five, where ...

AI infrastructure engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 100 K
throughput, and ensuring training is as efficient and cost-effective as possible. You'll also play a critical role in managing cluster orchestration with Slurm and Kubernetes while helping evolve the platform to support next-generation GPU infrastructure and specialised compute providers. This is an opportunity to work across ...

AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 80 K
will eliminate bottlenecks in the data path to ensure training is fast and as capital efficient as possible alongside managing cluster orchestration using slurm and Kubernetes while preparing to expand into specialised GPU providers. And finally you will master the stack from pytorch based learning libraries to complex data ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. Work with Slurm, Kubernetes, NVIDIA GPU Operator, NCCL and GPUDirect. Design high-performance storage solutions for AI workloads. Lead technical discussions with CTOs, AI leaders and infrastructure ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused computeBenefitsCompetitive salary and eligibility for equityBiannual bonus schemeFully expensed tech to match your needsPrivate health ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

Principal Cloud Architect – HPC/GPU & AI Platform Solutions

Hiring Organisation
Oracle Corporation
Location
London, United Kingdom
Salary
£ 80 K
public cloud platforms. Infrastructure as Code using Terraform and Ansible. Python, Bash, and PowerShell scripting. Kubernetes and container orchestration. HPC cluster management platforms including Slurm, PBS, or Bright Cluster Manager. High-performance networking technologies including RDMA and InfiniBand. MPI and distributed file systems. Cloud-native architectures. AI/… Large Language Models (LLMs) Agentic AI AI Platform Architecture Inference Serving Programming & AutomationPython Bash PowerShell Automation Frameworks Infrastructure Automation HPC TechnologiesSlurm PBS Bright Cluster Manager RDMA InfiniBand MPI Distributed File Systems Customer & ConsultingSolution Architecture Technical Consulting Pre-Sales Executive Presentations Customer Workshops Technical Enablement AI Transformation Strategy Cloud Adoption ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, United Kingdom
Salary
£ 80 K
scalable engineering solutions.Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows.Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability, and reliability … NVIDIA Nsight, DCGM, and related ecosystem technologies.Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation.Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks.Demonstrated success bridging research and production environments, enabling rapid experimentation while maintaining operational ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, UK
Employment Type
Full-time
engineering solutions. Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows. Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability … Nsight, DCGM, and related ecosystem technologies. Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation. Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks. Demonstrated success bridging research and production environments, enabling rapid experimentation while ...