1 to 25 of 250 OpenTelemetry Jobs in England

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands‐on technical authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical … transparency, innovation, and technical excellence that encourages continuous improvement and automation. Collaborative Enablement: Partner with development and SRE teams to drive the adoption of OpenTelemetry (OTel) and Google Cloud Observability (GCO) and Grafana standards. Regulatory Compliance: Operate effectively within a highly regulated environment, ensuring all observability and deployment solutions comply ...

Site Reliability Engineering Lead

Location
City Of London, England, United Kingdom
cloud‐native architectures. Expertise with Infrastructure as Code tools such as Terraform and Ansible. Strong background in observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, Datadog, or similar technologies. Experience managing large‐scale production environments with stringent availability requirements. Strong understanding of security, compliance, networking, and cloud governance principles. ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, UK
Employment Type
Full-time
telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your SkillsHands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support. Telemetry & Insights: Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR). Safe Production Automation: Develop proactive anomaly detection and Human … Hands‐on experience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks. Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems. Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/ ...

Software Engineer

Location
Greater London, England, United Kingdom
agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support. Telemetry & Insights: Implement ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR). Safe Production Automation: Develop proactive anomaly detection and Human … Systems: E xperience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks. Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems. Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/ ...

Senior Technical Consultant - DevOps

Hiring Organisation
CloudBees
Location
London, UK
Employment Type
Full-time
strategy, DevEx, or internal developer platforms (IDPs).Familiarity with AI-enabled development, agentic workflows, LLMs, or AI governance. Experience with observability/telemetry platforms (OpenTelemetry, Splunk, Dynatrace, Datadog, AppDynamics, Grafana).Experience in regulated or large-scale enterprise environments. Thought leadership via publications, conference speaking, or community involvement. Working Conditions: Travel ...

Sr. Systems Engineer

Hiring Organisation
Visa
Location
London, UK
Employment Type
Full-time
Experience building and maintaining CI/CD pipelines. Solid scripting experience using Python, Bash, or similar. Experience with monitoring/observability stacks (CloudWatch, Datadog, OpenTelemetry, etc.).Good understanding of networking fundamentals (DNS, routing, VPNs, load balancing, firewalls).Strong version-control discipline using GitHub or similar. Understanding of DevOps methodologies ...

Senior DevOps Engineer (Linux focused)

Hiring Organisation
Darktrace
Location
London, UK
Employment Type
Full-time
ArgoCD and Helm. You will also have strong experience of monitoring, logging, and observability practices, with expertise across technologies such as Prometheus, Grafana, Loki, OpenTelemetry, and the ELK stack. In addition, you will be able to demonstrate: Proven experience as a Senior DevOps Engineer, with a strong track record ...

Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

Hiring Organisation
FactSet Research Systems
Location
London, UK
Employment Type
Full-time
writtenundefinedAdditional Technical SkillsCloud Platforms:(e.g. AWS, GCP, Azure)CI/CD Tooling:(e.g. GitHub Actions, ArgoCD, Harness)Monitoring & Observability:(e.g. Prometheus, Grafana, Coralogix, OpenTelemetry)Infrastructure as Code:(e.g. Terraform, Pulumi)Config Management: (e.g. Ansible, Puppet, Chef)Programming/Scripting:(e.g. Python, Go, Bash)Soft Skills & General RequirementsStrong problem-solving ...

Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

Location
Greater London, England, United Kingdom
Additional Technical Skills Cloud Platforms: (e.g. AWS, GCP, Azure) CI/CD Tooling: (e.g. GitHub Actions, ArgoCD, Harness) Monitoring & Observability: (e.g. Prometheus, Grafana, Coralogix, OpenTelemetry) Infrastructure as Code: (e.g. Terraform, Pulumi) Config Management: (e.g. Ansible, Puppet, Chef) Programming/Scripting: (e.g. Python, Go, Bash) Soft Skills & General Requirements Strong problem ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Location
West Midlands, England, United Kingdom
specifically building and operating highly resilient cloud-native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
specifically building and operating highly resilient cloud‐native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
common cloud design patterns. Desirables Configuration management tooling (Puppet, Ansible). Advanced Kubernetes (Karpenter, KEDA, HPA/VPA, Service Mesh). Observability tooling (Datadog, OpenTelemetry) and SLI/SLO design. Security and identity tooling (SSO, IAM, PKI). Queue or streaming design patterns (SQS, Kafka). Commercial or low-level ...

Lead Software Engineer – DevSecOps & Platform Automation

Location
Greater London, England, United Kingdom
pipelines and infrastructure for cost, speed, and reliability; Observability, Reliability & Resilience: Integrate observability into delivery pipelines and platforms, including metrics, logs, and traces (e.g., OpenTelemetry), deployment health checks, and automated rollback signals; Implement and report on DORA and deployment reliability metrics (deployment frequency, lead time, change failure rate, MTTR); Support ...

Senior Site Reliability Engineer

Location
Cambridge, England, United Kingdom
technical strategy and roadmap decisions. Configuration management tooling (Puppet, Ansible). Advanced Kubernetes (Karpenter, KEDA, HPA/VPA, Service Mesh). Observability tooling (Datadog, OpenTelemetry) and mature SLI/SLO design. Security and identity tooling (SSO, IAM, PKI). Queue or streaming design patterns (SQS, Kafka). Resilience and disaster ...

Analytics Services Platform Engineer

Location
Greater London, England, United Kingdom
technologies including EMR, MSK, Athena, Redshift, Glue and MWAA Experience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetry Strong problem‐solving skills and a systematic approach to diagnosing and resolving issues Highly Desirable Skills Experience with streaming frameworks such as Flink, Kafka Streams ...

Database Platform Engineer

Location
Greater London, England, United Kingdom
services relevant to data platforms such as RDS, Aurora, S3, EC2 or EKS Familiarity with modern observability stacks such as Prometheus, Grafana, Elk or OTel Desirable: experience with cloud‐native and distributed SQL databases such as Aurora, YugabyteDB or TiDB Desirable: knowledge of data streaming and integration tools such ...

DevOps Engineer

Location
Greater London, England, United Kingdom
experience with policy‐as‐code frameworks for automated compliance and guardrails. Exposure to observability platforms such as Datadog, Prometheus/Grafana, or the OpenTelemetry ecosystem. Experience with container image hardening and scanning (Trivy, Grype, or similar). Experience with using AI tooling as well as a familiarity with security concerns ...

AI Platform Engineer- Senior Consultant-AI and Digital Factory

Location
Manchester, England, United Kingdom
Model Context Protocol (MCP)• Evaluation engineering (golden datasets, regression gates in CI, LLM-judge calibration)• Guardrail and AI-observability tooling (e.g. NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust)**MLOps & LLMOps**• Hands-on with MLOps platforms (Azure ML, Databricks, SageMaker) and vector/retrieval databases (Pinecone, Milvus, pgvector)• Experience with ...

AI Platform Engineer- Senior Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
Model Context Protocol (MCP)• Evaluation engineering (golden datasets, regression gates in CI, LLM-judge calibration)• Guardrail and AI-observability tooling (e.g. NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust)**MLOps & LLMOps**• Hands-on with MLOps platforms (Azure ML, Databricks, SageMaker) and vector/retrieval databases (Pinecone, Milvus, pgvector)• Experience with ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
roles, with leadership responsibilities. Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, AppDynamics, or Dynatrace. Hands-on experience with OpenTelemetry (OTel) for distributed tracing and observability instrumentation. Strong proficiency in Infrastructure as Code (IaC) using Terraform. Solid understanding of cloud platforms including ...

Head of Cloud

Location
Norwich, England, United Kingdom
expert), Kubernetes (EKS) IaC & Orchestration: Terraform, Helm, Terragrunt Languages: Go, Node.js, Python (automation/tooling) CI/CD: GitHub Actions, ArgoCD Observability: Prometheus, Grafana, OpenTelemetry What We’re Looking For Proven Leadership: Experience managing and scaling high-performing engineering teams. Cloud Expertise: Deep hands-on experience architecting and operating cloud ...

Platform Compute Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
storage, networking, IAM)Infrastructure-as-Code (e.g., Terraform, Pulumi)Monitoring and alerting (e.g., Prometheus, Grafana, Datadog, Zabbix)Logging and tracing (e.g., ELK stack, Fluentd, OpenTelemetry, JaegerIdentity and access management (e.g., LDAP, Kerberos, OAuth2)Cloud-native services (e.g., S3, EBS, GKE, EKS, Cloud Functions)Git-based workflows and version controlAgile methodologies ...

Site Reliability Engineer (SRE)

Location
Cambridge, England, United Kingdom
Software development experience(ideally working with and as a .NET developer) Strong understanding of SDLC, microservice and HA architecture Observability - NewRelic, ELK, Grafana, PagerDuty, OTEL or similar Experience with Kubernetes clusters in production setting, AWS, IOC Experience with operational tasks Knowledge of CI-CD tooling Jenkins, Gitlab, GitHub, ArgoCD ...