1 to 25 of 171 OpenTelemetry Jobs in England

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands‐on technical authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical … transparency, innovation, and technical excellence that encourages continuous improvement and automation. Collaborative Enablement: Partner with development and SRE teams to drive the adoption of OpenTelemetry (OTel) and Google Cloud Observability (GCO) and Grafana standards. Regulatory Compliance: Operate effectively within a highly regulated environment, ensuring all observability and deployment solutions comply ...

Site Reliability Engineering Lead

Location
City Of London, England, United Kingdom
cloud‐native architectures. Expertise with Infrastructure as Code tools such as Terraform and Ansible. Strong background in observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, Datadog, or similar technologies. Experience managing large‐scale production environments with stringent availability requirements. Strong understanding of security, compliance, networking, and cloud governance principles. ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support. Telemetry & Insights: Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR). Safe Production Automation: Develop proactive anomaly detection and Human … Hands‐on experience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks. Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems. Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/ ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Location
West Midlands, England, United Kingdom
specifically building and operating highly resilient cloud-native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software ...

Principal AI Platform Engineer (Python)

Location
Greater London, England, United Kingdom
test, package, and release software, and the practices around them, such as automated testing and staged rollouts. Observability: experience instrumenting systems with Prometheus, Grafana, OpenTelemetry, or an equivalent stack, and using that data to diagnose failures in distributed systems. Security (critical): a working grasp of secrets management, identity and access ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
specifically building and operating highly resilient cloud‐native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
common cloud design patterns. Desirables Configuration management tooling (Puppet, Ansible). Advanced Kubernetes (Karpenter, KEDA, HPA/VPA, Service Mesh). Observability tooling (Datadog, OpenTelemetry) and SLI/SLO design. Security and identity tooling (SSO, IAM, PKI). Queue or streaming design patterns (SQS, Kafka). Commercial or low-level ...

Senior Site Reliability Engineer

Location
Cambridge, England, United Kingdom
technical strategy and roadmap decisions. Configuration management tooling (Puppet, Ansible). Advanced Kubernetes (Karpenter, KEDA, HPA/VPA, Service Mesh). Observability tooling (Datadog, OpenTelemetry) and mature SLI/SLO design. Security and identity tooling (SSO, IAM, PKI). Queue or streaming design patterns (SQS, Kafka). Resilience and disaster ...

Cloud Native Specialist

Location
Greater London, England, United Kingdom
OpenShift) and container orchestration.* Demonstrate how Dynatrace provides automated, code-level visibility into microservices without manual instrumentation or sidecar overhead.* Advocate for OpenTelemetry (OTel) integration and explain how Dynatrace extends the value of open-source telemetry in a production-grade environment.2. Domain Execution:* Lead technical discovery and high-stakes Proof ...

Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

Location
Greater London, England, United Kingdom
Additional Technical Skills*** **Cloud Platforms:** *(e.g. AWS, GCP, Azure)** **CI/CD Tooling:** *(e.g. GitHub Actions, ArgoCD, Harness)** **Monitoring & Observability:** *(e.g. Prometheus, Grafana, Coralogix, OpenTelemetry)** **Infrastructure as Code:** *(e.g. Terraform, Pulumi)** **Config Management**: *(e.g. Ansible, Puppet, Chef)** **Programming/Scripting:** *(e.g. Python, Go, Bash)* **Soft Skills & General Requirements*** Strong problem ...

Analytics Services Platform Engineer

Location
Greater London, England, United Kingdom
technologies including EMR, MSK, Athena, Redshift, Glue and MWAA Experience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetry Strong problem‐solving skills and a systematic approach to diagnosing and resolving issues Highly Desirable Skills Experience with streaming frameworks such as Flink, Kafka Streams ...

Head of Cloud

Location
Norwich, England, United Kingdom
expert), Kubernetes (EKS) IaC & Orchestration: Terraform, Helm, Terragrunt Languages: Go, Node.js, Python (automation/tooling) CI/CD: GitHub Actions, ArgoCD Observability: Prometheus, Grafana, OpenTelemetry What We’re Looking For Proven Leadership: Experience managing and scaling high-performing engineering teams. Cloud Expertise: Deep hands-on experience architecting and operating cloud ...

Site Reliability Engineer - Service Assurance Systems

Location
Greater London, England, United Kingdom
systems at scale. Familiarity with infrastructure-as-code tools such as Terraform or Ansible. Experience with log aggregation and analysis platforms such as the OTEL Stack or AWS CloudWatch Logs Insights. Exposure to Kubernetes or other container orchestration platforms. Experience working in an Agile or DevOps team environment. EEO Statement ...

Site Reliability Engineer

Location
Newcastle upon Tyne, England, United Kingdom
Cloudwatch, Sumologic etc. Extensive understanding of networking and security concepts. Bonus Points For Specialized SRE observability experience with New Relic or DataDog. Familiarity with OpenTelemetry, AIOps, MLOps, or SecOps. Location: Newcastle, UK - In-Office (at least 4 days per week at the office) Why You'll Love Working With ...

DevOps Engineer

Location
Greater London, England, United Kingdom
experience with policy‐as‐code frameworks for automated compliance and guardrails. Exposure to observability platforms such as Datadog, Prometheus/Grafana, or the OpenTelemetry ecosystem. Experience with container image hardening and scanning (Trivy, Grype, or similar). Experience with using AI tooling as well as a familiarity with security concerns ...

Logs Specialist

Location
Maidenhead, England, United Kingdom
powered observability platform.**Core Responsibilities****Domain Expertise:*** Act as the domain "subject matter expert" (SME) for Logs, staying ahead of industry trends like OpenTelemetry (OTel), log pipelines (Cribl/BindPane), and cloud-native logging (CloudWatch/Stackdriver).* Articulate the architectural superiority of Dynatrace Grail—specifically how its schema … optimize their SaaS consumption and maximize their Dynatrace investment.**Solution Architecture:*** Expertise in architecting resilient, vendor‐neutral log‐ingestion frameworks utilizing Fluentd, Logstash, and OpenTelemetry Collector pipelines, etc.* Help customers navigate complex log-routing scenarios, ensuring high-value data is prioritized for analytics while low-value data is archived cost ...

Senior Software Development Engineer (SRE)

Location
Cambridge, England, United Kingdom
Software development experience(ideally working with and as a .NET developer) Strong understanding of SDLC, microservice and HA architecture Observability - NewRelic, ELK, Grafana, PagerDuty, OTEL or similar Experience with Kubernetes clusters in production setting, AWS, IOC Experience with operational tasks Knowledge of CI-CD tooling Jenkins, Gitlab, GitHub, ArgoCD ...

Database Platform Engineer

Location
Greater London, England, United Kingdom
services relevant to data platforms such as RDS, Aurora, S3, EC2 or EKS Familiarity with modern observability stacks such as Prometheus, Grafana, Elk or OTel Desirable: experience with cloud‐native and distributed SQL databases such as Aurora, YugabyteDB or TiDB Desirable: knowledge of data streaming and integration tools such ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop, self-healing ...

Senior Software Engineer

Location
Reading, England, United Kingdom
Kubernetes and cloud‐native deployments CI/CD pipeline experience (GitLab, GitHub or similar) Infrastructure as Code (Terraform or similar) Experience with observability tooling (OpenTelemetry, Prometheus, Grafana, etc.) Strong testing mindset (TDD, automated testing, contract testing) Highly Desirable Experience with API Gateway technologies (e.g. AWS API Gateway) Experience building ...

Principal DevOps Engineer

Location
Nottingham, England, United Kingdom
KEDA or Karpenter Falco or Kyverno Go development Azure Front Door Azure Firewall Entra ID and PIM Observability platforms such as Grafana, Loki or OpenTelemetry Experience working within regulated or compliance-focused environments #J-18808-Ljbffr ...

Camunda Architect / Lead Architect (Camunda 8 Preferred)

Location
Greater London, England, United Kingdom
SaaS and self-managed deployments at scale Exposure to cloud platforms such as: AWS Azure GCP Experience with observability tooling such as: Prometheus Grafana OpenTelemetry Experience with: enterprise integration patterns API management workflow/task UI customisation Experience in regulated industries such as: Banking Insurance Healthcare Ideal Profile This role ...

Lead Software Engineer

Location
Greater London, England, United Kingdom
customer support requirements. Experience with Kafka and event-driven architectures. Experience with OAuth2, OpenID Connect, SAML, and Keycloak. Experience with Datadog, Grafana, Elastic, OpenTelemetry, or similar observability tooling. Experience implementing AI-assisted software development practices. WSD is an employer that values diversity.We highly encourage applications from appropriately qualified and eligible ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
pipelines using GitHub Actions or GitLab CI/CD, plus ArgoCD (GitOps). Experience with monitoring/observability tools (Prometheus, Grafana, ELK stack, Datadog, OpenTelemetry) – including metrics, logs, traces, dashboards, alerting, and integration with AWS services (e.g., CloudWatch, X‐Ray). Solid understanding of AWS EKS IAM concepts, especially ...