1 to 25 of 90 Slurm Workload Manager Jobs in the UK

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Machine Learning Engineer (Mid to Principal)

Location
West of England, England, United Kingdom
Jira. Edge deployments: Nvidia Jetson (e.g. AGX Orin), Raspberry Pi, or other embedded accelerators. Distributed model training & infra: Pytorch DDP, FDSP and TorchTitan, Megatron, Slurm, Run:ai, DeepSpeed, Kubernetes, cloud or on‐prem GPU clusters. About you You’ve built ML systems that persist—deployed in real settings, iterated ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 100 K
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub.Knowledge of other programming ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub. Knowledge of other ...

Infrastructure Engineer (Linux)

Location
Greater London, England, United Kingdom
storage issues (network filesystems, capacity management, performance troubleshooting). Support day-to-day operation of a cluster compute/job scheduling environment (e.g. Slurm). Help with user and identity management tasks (e.g. FreeIPA/LDAP), including onboarding and access provisioning. Document infrastructure changes and contribute to runbooks ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 70 K
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

HPC System Administrator

Location
Farnborough, England, United Kingdom
defects. Perform Software and firmware testing for any fixes, upgrades, security patch. Position requirements: Hands-on experience with HPC tools and schedulers such as Slurm and PBS Pro, as well as Confluent Solid scripting and automation skills using Ansible, Bash, and Python Experience with distributed storage systems and file ...

Senior Software Engineer - Scientific / HPC

Hiring Organisation
Technical Futures Ltd
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 - £90,000 per annum
code quality. Some/most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible ...

Senior HPC Systems Administrator

Location
Oxford, England, United Kingdom
systems integration skills Excellent written and verbal communication abilities Ability to work independently and collaboratively in a research environment Desirable Skills Experience with SLURM or other job scheduling systems Experience with containerisation and virtualisation technologies Knowledge of cloud computing platforms Familiarity with scientific computing tools and AI/ ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads.AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms.ML frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent … ArgoCD) for continuous, declarative platform delivery.Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines.Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms. ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron … continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

HPC System Administrator — On-Site/Remote Linux & HPC Ops

Location
Farnborough, England, United Kingdom
System Administrator in London, UK, to manage data center infrastructure, work across on-site and remote support, and ensure high availability. The role emphasizes Slurm and PBS Pro, Linux administration, and hands-on hardware familiarity within a Managed Services context. The team seeks experience with HPC tooling, scripting (Python ...

Senior HPC Linux Engineer — On‐Site, Security‐Clearance Ready

Location
West of England, England, United Kingdom
system security and performance monitoring. The role requires a 3+ year Linux background, scripting skills (Bash/Python), and a willingness to learn Slurm, xCAT and Ansible. On-site work five days a week with UK National status and clearance eligibility. #J-18808-Ljbffr ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
titles matter less than experience. We're particularly interested in people with experience in:Large-scale distributed training.PyTorch and modern deep learning frameworks.Kubernetes, Slurm or GPU orchestration platforms.AWS and specialist GPU cloud providers.High-performance computing and distributed systems.Training optimisation, memory management and networking.MLOps tooling, CI/CD and experimentation ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
matter less than experience. We're particularly interested in people with experience in: Large-scale distributed training. PyTorch and modern deep learning frameworks. Kubernetes, Slurm or GPU orchestration platforms. AWS and specialist GPU cloud providers. High-performance computing and distributed systems. Training optimisation, memory management and networking. MLOps tooling ...

Senior HPC Systems Administrator

Hiring Organisation
University of Oxford
Location
Oxford, Oxfordshire, United Kingdom
Salary
£ 50 K
troubleshooting and systems integration skills● Excellent written and verbal communication abilities● Ability to work independently and collaboratively in a research environmentDesirable Skills● Experience with SLURM or other job scheduling systems● Experience with containerisation and virtualisation technologies● Knowledge of cloud computing platforms● Familiarity with scientific computing tools and AI/ ...

HPC Systems Engineer

Location
Greater London, England, United Kingdom
Need: At least 5+ years of professional experience in high performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required At least 5+ years of experience with Linux systems administration High proficiency ...

Senior Applied Research Engineer - Video Team Synthesia Europe, Germany, Switzerland, UK

Location
United Kingdom
models Experience with GANs or VAEs Experience optimizing inference systems for production Our stack Python, PyTorch, CUDA DeepSpeed, distributed training & inference Sequence parallelism AWS, SLURM, Docker GitHub, CI/CD pipelines Who you are You are research-driven but outcome-focused You care about shipping, not just publishing ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, United Kingdom
Salary
£ 80 K
scale distributed training or HPC systems.Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.Experience working with containerized environments (Docker, Singularity/ ...

Infrastructure Engineer

Location
Cambridge, England, United Kingdom
LDAP, NIS and DHCP. Experience with NFS/enterprise storage. FlexLM experience is desirable. Experience with Ansible, Puppet, Chef or Salt. Knowledge of Slurm, Grid Engine, LSF or similar is advantageous. #J-18808-Ljbffr ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Senior kdb+ Developer, Vice President

Hiring Organisation
State Street Bank
Location
London, United Kingdom
Salary
£ 80 K
Python.Proficiency with modern data science toolchains, including Jupyter, pandas, NumPy, scikit‐learn, with applied machine‐learning experience.Solid understanding of parallel computing frameworks such as Slurm or equivalent technologies.Strong background in quantitative analysis, including mathematical modelling, statistics, regression, and probability theory.Bachelor’s or Master’s degree in Computer Science, Mathematics ...