17 of 17 Chaos Engineering Jobs in London

DevOps Consultant | Harness London

Hiring Organisation
Infosys Technologies
Location
London, United Kingdom
Salary
£ 70 K
seeking a seasoned Harness Platform Specialist with deep expertise across the Harness Software Delivery Platform, including Continuous Integration, Continuous Delivery, Security Testing Orchestration (STO), Chaos Engineering, Cloud Cost Management (CCM), and Harness AI Agents. The ideal candidate will bring hands on experience in designing scalable CI/… Harness platform, aligned with enterprise delivery, security, and compliance standards.Architect and maintain end‐to‐end delivery workflows across Harness modules, including CI, CD, STO, Chaos Engineering, CCM, and emerging AI‐driven capabilities.Implement and manage Harness Delegate installations, upgrades, image customizations, and integrations across cloud and on‐prem environments.Collaborate ...

SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Location
Greater London, England, United Kingdom
passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving the transformation from reactive operations … proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce non-functional ...

SRE Architect (68019)

Hiring Organisation
Hitachi
Location
London, United Kingdom
Salary
£ 80 K
perspective, and passion for achieving great things in the world are equally as important to us.Job descriptionMandatory Skills:Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management.Role Description Skills:ROLE PURPOSELead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive … engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving.KEY RESPONSIBILITIES• Define and enforce non-functional requirements (NFRs ...

Lead Site Reliability Engineer (AWS)

Location
Greater London, England, United Kingdom
excellence of the the company's global digital platforms. With a strong specialism in AWS, along with demonstrable experience, you will work closely with engineering teams, architects, senior stakeholders, and our managed service partner, you will design, implement, and continuously improve resilient, scalable, and secure systems. You will leverage … incident management to deliver high system uptime and support evolving business needs. About the Team The role sits within the Digital and Technology Engineering team, which partners across the the company to deliver customer‐centric digital products and services. The Engineering function brings together Architecture, Software Engineering ...

VodafoneThree - SRE III

Hiring Organisation
Vodafone
Location
London, United Kingdom
Salary
£ 80 K
innovative teams, creating a connected future with technologies like Cloud, AI and big data.What you’ll doIn this role, within VodafoneThree's Performance and Chaos Engineering (PaCE) team, you will play a key role in enabling engineering teams to deliver scalable, resilient, and high-performing digital services. … teams identify and address risks early and build confidence in their solutions before they reach production.Working closely with product, platform, development, and Site Reliability Engineering teams, you will champion a shift-left approach to performance and resilience engineering. Your focus will be on providing the frameworks, tooling, standards ...

VodafoneThree - SRE III

Hiring Organisation
VodafoneThree
Location
Greater London, United Kingdom
Employment Type
Full Time
creating a connected future with technologies like Cloud, AI and big data. What you'll do In this role, within VodafoneThree's Performance and Chaos Engineering (PaCE) team, you will play a key role in enabling engineering teams to deliver scalable, resilient, and high-performing digital services. … identify and address risks early and build confidence in their solutions before they reach production. Working closely with product, platform, development, and Site Reliability Engineering teams, you will champion a shift-left approach to performance and resilience engineering. Your focus will be on providing the frameworks, tooling, standards ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Source Group International
Location
London, United Kingdom
Salary
£ 80 K
will define reliability in measurable terms and build the tooling and processes to achieve it, improving platform speed, stability, and scalability.Key responsibilities Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions.Enhance observability across services (metrics, logs, traces) to detect … right-size workloads, tune autoscaling, and improve infrastructure efficiency.Improve production readiness via pre-deployment checks, post-release validation, and robust platform guardrails.Introduce and run chaos engineering experiments to strengthen resilience and recovery.Automate operational processes to reduce manual intervention and toil across the stack.Support major incident response, root-cause ...

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
performance bottlenecks, optimize resource utilization, and guide capacity planning.* Lead & Mentor: Act as a technical leader and mentor for the observability team and wider engineering groups. Champion and enforce best practices, fostering a culture of proactive and data-informed decision-making.* Drive Incident & Problem Management: Working with Operations teams … part of this.**Job Requirements:**Essential Qualifications* Experience: 5-7+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a proven track record of leading complex projects.* Technical Leadership: Demonstrated experience in architecting and designing large-scale monitoring and observability solutions. ...

Architect & Delivery Lead (68018)

Hiring Organisation
Hitachi
Location
London, United Kingdom
Salary
£ 120 K
Make architectural decisions across IAM, Cloud, SRE, Network, Data, and Security — ensuring coherence, reusability, and alignment with business objectives• Establish and chair the Architecture & Engineering Governance board, providing technical assurance across all workstreams• Own the programme roadmap, resource plan, and financial model — tracking cost savings, team reduction trajectory … vendor and tool selection, ensuring standardisation across the programme and eliminating redundant tooling• Build and lead high-performing distributed teams, fostering a culture of engineering excellence, accountability, and continuous improvement• Define the continuous improvement factory model, ensuring the transformation sustains beyond the initial programmeTECHNICAL SKILLS & EXPERTISE• Broad and deep ...

Architect & Delivery Lead (68018)

Location
Greater London, England, United Kingdom
Make architectural decisions across IAM, Cloud, SRE, Network, Data, and Security — ensuring coherence, reusability, and alignment with business objectives Establish and chair the Architecture & Engineering Governance board, providing technical assurance across all workstreams Own the programme roadmap, resource plan, and financial model — tracking cost savings, team reduction trajectory … vendor and tool selection, ensuring standardisation across the programme and eliminating redundant tooling Build and lead high‐performing distributed teams, fostering a culture of engineering excellence, accountability, and continuous improvement Define the continuous improvement factory model, ensuring the transformation sustains beyond the initial programme Technical Skills & Expertise Broad ...

Architect & Delivery Lead (68018)

Location
Greater London, England, United Kingdom
dependenciesMake architectural decisions across IAM, Cloud, SRE, Network, Data, and Security — ensuring coherence, reusability, and alignment with business objectivesEstablish and chair the Architecture & Engineering Governance board, providing technical assurance across all workstreamsOwn the programme roadmap, resource plan, and financial model — tracking cost savings, team reduction trajectory, and ROIDrive stakeholder … stakeholdersDrive vendor and tool selection, ensuring standardisation across the programme and eliminating redundant toolingBuild and lead high-performing distributed teams, fostering a culture of engineering excellence, accountability, and continuous improvementDefine the continuous improvement factory model, ensuring the transformation sustains beyond the initial programmeTECHNICAL SKILLS & EXPERTISEBroad and deep technical knowledge ...

Lead DevOps Engineer

Hiring Organisation
Elliptic
Location
London, United Kingdom
Salary
£ 80 K
hands-on team Lead role where you will balance technical expertise with team leadership.Elliptic is building an AI fluent workforce. Across our Product, Engineering, Design and Intelligence we’re going beyond giving everyone the tools. We’re setting new standards. Using AI to transform how we build and work … from day one.What you will do:Own the DevOps and Platform roadmap, including Kubernetes platform evolution, application packaging and migration to EKS, and enabling engineering teams to ship reliably to production.Lead by doing - engineer, review, and enhance Kubernetes and CNCF-aligned infrastructure, setting technical standards.Architect multi-cluster, multi-region ...

Engineering Manager - Observability (Hybrid, London)

Hiring Organisation
CrowdStrike
Location
London, United Kingdom
Salary
£ 70 K
community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.About the Role:CrowdStrike is hiring an Engineering Manager - Observability to help take our metrics and tracing to the next level. We are looking for a highly-technical engineering leader with … teams and experience with open source projects commonly found in large-scale deployments. Our team works to develop infrastructure services to support the CrowdStrike engineering teams’ pursuit of a full DevOps model.This is a hybrid role based in one of our offices inLondon (United Kingdom), Aarhus (Denmark) or Dublin ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
resiliency patterns to enhance application reliability. Recovery Testing Support: Support and participate in advanced recovery testing, including Production Swing Tests, Data Recovery Tests, and chaos engineering practices. Automation Drive: Drive the adoption and development of automation solutions (such as Ansible playbooks and Terraform) to minimize recovery time … application teams, SRE leads, and stakeholders. Qualifications Significant professional experience in software development, or an equivalent field, with a strong focus on Site Reliability Engineering and Observability. Expertise in analyzing complex application, database, network, and OS issues within large-scale, customer‐facing systems. A service‐oriented attitude combined with ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
resiliency patterns to enhance application reliability. Recovery Testing Support: Support and participate in advanced recovery testing, including Production Swing Tests, Data Recovery Tests, and chaos engineering practices. Automation Drive: Drive the adoption and development of automation solutions (such as Ansible playbooks and Terraform) to minimize recovery time … application teams, SRE leads, and stakeholders. Qualifications Significant professional experience in software development, or an equivalent field, with a strong focus on Site Reliability Engineering and Observability. Expertise in analyzing complex application, database, network, and OS issues within large-scale, customer-facing systems. A service-oriented attitude combined with ...

Staff Software Engineer, AI Reliability Engineering

Hiring Organisation
Humanloop
Location
London, United Kingdom
Salary
> £ 150 K
systems.About the RoleClaude has your back. AIRE has Claude's. Help us keep Claude reliable for everyone who depends on it.AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API layers, serving … hardware accelerators (GPUs, TPUs, Trainium).Understand ML-specific networking optimizations like RDMA and InfiniBand.Have expertise in AI-specific observability tools and frameworks.Have experience with chaos engineering and systematic resilience testing.Have contributed to open-source infrastructure or ML tooling.The annual compensation range for this role is listed below. ...

Senior Backend Engineer: Chaos & Reliability (Remote)

Location
Greater London, England, United Kingdom
Camunda seeks a Senior Software Engineer, Backend, to own automated reliability testing and chaos engineering for Camunda 8. You will break things in safe environments to strengthen the platform, guide product direction, and collaborate with friendly colleagues who live our FAITH values. You’ll design and run chaos ...