1 to 25 of 128 OpenTelemetry Jobs in England

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands-on technical authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical … transparency, innovation, and technical excellence that encourages continuous improvement and automation. Collaborative Enablement: Partner with development and SRE teams to drive the adoption of OpenTelemetry (OTel) and Google Cloud Observability (GCO) and Grafana standards. Regulatory Compliance: Operate effectively within a highly regulated environment, ensuring all observability and deployment solutions comply ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated ...

Strategic DevSecOps Consultant

Hiring Organisation
CloudBees
Location
London, UK
Employment Type
Full-time
.Familiarity with AI-enabled software development, agentic workflows, large language models (LLMs), or AI governance practices. Experience with observability and telemetry platforms such as OpenTelemetry, Splunk, Dynatrace, Datadog, AppDynamics, Grafana, or similar technologies. Experience working with large-scale enterprise architecture, governance, compliance, and regulated environments. Thought leadership experience through technical ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support. Telemetry & Insights: Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR). Safe Production Automation: Develop proactive anomaly detection and Human … Hands‐on experience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks. Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems. Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/ ...

Principal AI Platform Engineer (Python)

Location
Greater London, England, United Kingdom
test, package, and release software, and the practices around them, such as automated testing and staged rollouts. Observability: experience instrumenting systems with Prometheus, Grafana, OpenTelemetry, or an equivalent stack, and using that data to diagnose failures in distributed systems. Security (critical): a working grasp of secrets management, identity and access ...

Cloud Native Specialist

Location
Greater London, England, United Kingdom
OpenShift) and container orchestration.* Demonstrate how Dynatrace provides automated, code-level visibility into microservices without manual instrumentation or sidecar overhead.* Advocate for OpenTelemetry (OTel) integration and explain how Dynatrace extends the value of open-source telemetry in a production-grade environment.2. Domain Execution:* Lead technical discovery and high-stakes Proof ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Location
Birmingham, England, United Kingdom
specifically building and operating highly resilient cloud-native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
specifically building and operating highly resilient cloud‐native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software ...

Senior Cloud Engineer, AI Platform SRE

Location
Manchester, England, United Kingdom
Built and maintained CI/CD with GitHub Actions, GitLab CI, Argo CD, Jenkins or similar. Observability: Hands‐on with Datadog, Prometheus, Grafana or OpenTelemetry, and opinionated about what's worth alerting on. Incident management: Calm, methodical instincts under pressure, and a habit of fixing the class of problem rather ...

Senior Cloud Engineer, AI Platform SRE

Location
Leeds, England, United Kingdom
Built and maintained CI/CD with GitHub Actions, GitLab CI, Argo CD, Jenkins or similar. Observability: Hands‐on with Datadog, Prometheus, Grafana or OpenTelemetry, and opinionated about what's worth alerting on. Incident management: Calm, methodical instincts under pressure, and a habit of fixing the class of problem rather ...

Senior Cloud Engineer, AI Platform SRE

Location
Greater London, England, United Kingdom
Built and maintained CI/CD with GitHub Actions, GitLab CI, Argo CD, Jenkins or similar. Observability: Hands‐on with Datadog, Prometheus, Grafana or OpenTelemetry, and opinionated about what's worth alerting on. Incident management: Calm, methodical instincts under pressure, and a habit of fixing the class of problem rather ...

Head of Cloud

Location
Norwich, England, United Kingdom
expert), Kubernetes (EKS) IaC & Orchestration: Terraform, Helm, Terragrunt Languages: Go, Node.js, Python (automation/tooling) CI/CD: GitHub Actions, ArgoCD Observability: Prometheus, Grafana, OpenTelemetry What We’re Looking For Proven Leadership: Experience managing and scaling high-performing engineering teams. Cloud Expertise: Deep hands-on experience architecting and operating cloud ...

Site Reliability Engineer - Service Assurance Systems

Location
Greater London, England, United Kingdom
systems at scale. Familiarity with infrastructure-as-code tools such as Terraform or Ansible. Experience with log aggregation and analysis platforms such as the OTEL Stack or AWS CloudWatch Logs Insights. Exposure to Kubernetes or other container orchestration platforms. Experience working in an Agile or DevOps team environment. EEO Statement ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
roles, with leadership responsibilities. Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, AppDynamics, or Dynatrace. Hands-on experience with OpenTelemetry (OTel) for distributed tracing and observability instrumentation. Strong proficiency in Infrastructure as Code (IaC) using Terraform. Solid understanding of cloud platforms including ...

Platform Compute Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
storage, networking, IAM)Infrastructure-as-Code (e.g., Terraform, Pulumi)Monitoring and alerting (e.g., Prometheus, Grafana, Datadog, Zabbix)Logging and tracing (e.g., ELK stack, Fluentd, OpenTelemetry, JaegerIdentity and access management (e.g., LDAP, Kerberos, OAuth2)Cloud-native services (e.g., S3, EBS, GKE, EKS, Cloud Functions)Git-based workflows and version controlAgile methodologies ...

Staff Engineer - Money, Risk & Payment Ancillaries Yuno Totalmente remoto · Mundial ayer

Location
Greater London, England, United Kingdom
gRPC, REST Frameworks — Spring Boot, Spring WebFlux; Go standard library Messaging — Apache Kafka, SQS Databases — PostgreSQL, Redis Infrastructure — AWS, Kubernetes, Docker, Terraform Observability — Datadog, OpenTelemetry CI/CD — GitHub Actions, ArgoCD Version Control — Git/GitHub What We Offer at Yuno Competitive Compensation Remote Work — you can work from everywhere ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop, self-healing ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
secure and fit for purpose. Technology Environment AWS Terraform Docker Kafka/AWS MSK Couchbase MongoDB Redis ClickHouse GitHub Actions/GitLab CI Grafana OpenTelemetry Python Golang What We're Looking For Strong experience building and operating AWS-based infrastructure platforms. Expert-level Terraform experience within enterprise production environments. Extensive ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
latency, cost, hallucination/quality metrics) using tools such as Arize, WhyLabs, or Langfuse. Logging & tracing: centralized logging (ELK/OpenSearch) and distributed tracing (OpenTelemetry) across data, training, and inference pipelines for end-to-end root-cause analysis. Integration — AI Stack, Enterprise Networks & Service Provider EnvironmentsPlatform integration: API-based ...

Senior Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
City of London, London, United Kingdom
failover, and production-recovery exercises, highlighting skills in system reliability and continuity planning. Experience with enterprise observability tools such as Splunk, ELK, Grafana, Prometheus, OpenTelemetry, AppDynamics, or Dynatrace, reflecting proficiency in monitoring and diagnostics. Experience modernizing monolithic or legacy enterprise applications into maintain ...

Tech Lead, Site Reliability Engineering

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, UK
Employment Type
Full-time
orchestration (Kubernetes, Docker).Proficiency in CI/CD pipelines and infrastructure-as-code tools (Terraform, GitHub Actions, Jenkins).Familiarity with observability platforms (Datadog, BigPanda, OpenTelemetry).Experience working with identity platforms and/or fraud detection systems. Excellent communication and stakeholder management skills; ability to influence across technical and business domains. ...

Engineer C# (Full Stack)

Location
City Of London, England, United Kingdom
with microservices and event‐driven architectures. Experience with AI‐assisted design and coding. Knowledge of GraphQL and WebSockets. Familiarity with observability tools such as OpenTelemetry and Grafana. Experience with Infrastructure as Code (Terraform). Understanding of financial markets or trading systems. Contribution to open‐source projects. Awareness of security principles ...

QA Engineer

Hiring Organisation
Moorepay
Location
Manchester, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
with performance and load testing tools (e.g. JMeter, k6, Locust). Understanding of contract testing (e.g., Pact). Familiarity with observability tools (e.g., Grafana, OpenTelemetry). ...

Senior DevOps Engineer

Hiring Organisation
Hackajob Ltd
Location
Leicester, Leicestershire, East Midlands, United Kingdom
Employment Type
Permanent
Salary
£70,000
scripting Networking skills/fundamentals? SQL Server/NoSQL databases Windows & Linux servers Source Control Management (Git) Docker containers Kubernetes Microservices Monitoring tool - Dynatrace, OTel, Grafana Docker Compose/Helm charts DevSecOps Tooling - SonarCloud/PrismaCloud/CrowdStrike ...

Engineer C# (Full Stack)

Location
Greater London, England, United Kingdom
collaboration skills.Desired* Experience with microservices and event-driven architectures.* Experience with AI assisted design and coding.* GraphQL, and WebSockets.* Knowledge of observability tools (e.g., OpenTelemetry, Grafana).* Familiarity with Infrastructure as Code (Terraform).* Understanding of financial markets or trading systems.* Contribution to open-source projects.* Awareness of security principles ...