1 to 25 of 72 Slurm Workload Manager Jobs in the UK

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Machine Learning Engineer (Mid to Principal)

Location
West of England, England, United Kingdom
Jira. Edge deployments: Nvidia Jetson (e.g. AGX Orin), Raspberry Pi, or other embedded accelerators. Distributed model training & infra: Pytorch DDP, FDSP and TorchTitan, Megatron, Slurm, Run:ai, DeepSpeed, Kubernetes, cloud or on‐prem GPU clusters. About you You’ve built ML systems that persist—deployed in real settings, iterated ...

Platform Application Specialist

Hiring Organisation
Appcast
Location
London, UK
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub.Knowledge of other programming ...

Infrastructure Engineer (Linux)

Location
Greater London, England, United Kingdom
storage issues (network filesystems, capacity management, performance troubleshooting). Support day-to-day operation of a cluster compute/job scheduling environment (e.g. Slurm). Help with user and identity management tasks (e.g. FreeIPA/LDAP), including onboarding and access provisioning. Document infrastructure changes and contribute to runbooks ...

Junior Platform Specialist

Hiring Organisation
Appcast
Location
London, UK
Gitlab, Artifactory or Docker,Experience with infrastructure automation and configuration management, such as Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

HPC System Administrator

Location
Farnborough, England, United Kingdom
defects. Perform Software and firmware testing for any fixes, upgrades, security patch. Position requirements: Hands-on experience with HPC tools and schedulers such as Slurm and PBS Pro, as well as Confluent Solid scripting and automation skills using Ansible, Bash, and Python Experience with distributed storage systems and file ...

Senior Software Engineer - Scientific / HPC

Hiring Organisation
Technical Futures Ltd
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 - £90,000 per annum
code quality. Some/most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms. ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron … continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

Senior HPC Linux Engineer — On‐Site, Security‐Clearance Ready

Location
West of England, England, United Kingdom
system security and performance monitoring. The role requires a 3+ year Linux background, scripting skills (Bash/Python), and a willingness to learn Slurm, xCAT and Ansible. On-site work five days a week with UK National status and clearance eligibility. #J-18808-Ljbffr ...

Founding AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
matter less than experience. We're particularly interested in people with experience in: Large-scale distributed training. PyTorch and modern deep learning frameworks. Kubernetes, Slurm or GPU orchestration platforms. AWS and specialist GPU cloud providers. High-performance computing and distributed systems. Training optimisation, memory management and networking. MLOps tooling ...

HPC Systems Engineer

Location
Greater London, England, United Kingdom
Need: At least 5+ years of professional experience in high performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required At least 5+ years of experience with Linux systems administration High proficiency ...

Senior Applied Research Engineer - Video Team Synthesia Europe, Germany, Switzerland, UK

Location
United Kingdom
models Experience with GANs or VAEs Experience optimizing inference systems for production Our stack Python, PyTorch, CUDA DeepSpeed, distributed training & inference Sequence parallelism AWS, SLURM, Docker GitHub, CI/CD pipelines Who you are You are research-driven but outcome-focused You care about shipping, not just publishing ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker ...

Hybrid HPC System Administrator – Slurm, Linux & Storage

Location
Farnborough, England, United Kingdom
Lenovo in the United Kingdom (Hampshire) is hiring an HPC System Administrator to manage data center infrastructure, Linux/Unix environments, and monitoring to ensure uptime and security. This role blends on-site and remote ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Senior kdb+ Developer, Vice President

Hiring Organisation
State Street Bank
Location
London, UK
Employment Type
Full-time
with modern data science toolchains, including Jupyter, pandas, NumPy, scikit‐learn, with applied machine‐learning experience. Solid understanding of parallel computing frameworks such as Slurm or equivalent technologies. Strong background in quantitative analysis, including mathematical modelling, statistics, regression, and probability theory. Bachelor's or Master's degree in Computer ...

Senior Data & MLOps Engineer

Location
Greater London, England, United Kingdom
least one systems language (Go, Rust, or C++). Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm). Familiarity with GPU telemetry systems such as NVML or DCGM and hardware‐level monitoring concepts. Demonstrated experience scaling systems from Proof‐of‐Concept ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, UK
Employment Type
Full-time
cost-efficiency. Your ImpactProvision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed networking … from individuals. Anaplan does not: Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person. Send job offers via email. All offers are first extended verbally by a member of our internal recruitment team ...

Senior Field Engineer - AI/ML HPC & Kubernetes Solutions

Location
Greater London, England, United Kingdom
drive proofs of concept to accelerate client deployments. You will collaborate with engineering teams, contribute to product direction, and help optimize workloads using Slurm, NCCL, and Infiniband in high-performance environments. #J-18808-Ljbffr ...

Lead AI Infrastructure & Distributed Systems Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
engineer who thrives in early stage startup environments and prefers broad systems ownership over narrow specialisation. Technical Expertise: Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks. Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking is essential. Relevant Background: Experience ...

Research Engineer, Forge

Location
Greater London, England, United Kingdom
comfort in fast‐moving, under‐specified environments. Nice to have Distributed training experience (FSDP, DeepSpeed, Megatron, etc.). Cluster/orchestration experience (SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, etc.). Experience building reliable ML infrastructure, evaluation systems, or large‐scale data processing pipelines. Research experience in LLMs, agents, multimodal ...

Bioinformatician Plant Biology Institute Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
bioinformatics tools for genome assembly, annotation and variant analysis. Familiarity operating within Linux environments, including operating within HPC and/or cloud environments (e.g. SLURM, OCI). Experience applying machine learning methods to genomic and transcriptomic data, with an understanding of different ML approaches and their suitability for different ...

Infrastructure Engineer (Cambridge)

Location
Cambridge, England, United Kingdom
/reliable systems, Care about the experience of the people relying on your infrastructure. Nice to have/Beneficial HPC‐style job scheduling (Kueue, Slurm, MPI, or similar). Experience with large scientific or array datasets. Any background in photonics, semiconductors, EDA, or scientific computing. We are particularly interested ...