CareerPlanGet AI match score →

Senior Platform Engineer

🌐 Remote💼 Full-time🗓 2026-06-10 → 2026-07-27

Required skills

Demonstrable, hands-on experience applying core DevOps and Site Reliability Engineering (SRE) principles to manage, monitor, and scale production systems, A deep understanding of the SRE mindset, including SLO/SLA creation and monitoring, error budget management, toil reduction, and post-incident review (blameless postmortems), Expertise in one or more major public cloud providers (AWS, GCP, or Azure), encompassing network configuration, security best practices (IAM, security groups, etc.), compute services (EC2, GKE, ECS, etc.), and managed services (databases, queues, serverless functions), In-depth knowledge of container technologies, specifically Docker, and extensive experience orchestrating them at scale using Kubernetes (K8s), Proficiency in one or more modern software languages (e.g., Typescript, Go, Python, Rust) and associated frameworks used for building high-performance, resilient production systems, Proven experience developing robust, maintainable, and well-tested automation scripts, services and pipelines to manage infrastructure, deployments, and operational tasks, Experience owning, managing, and maintaining mission-critical operational tooling

Preferred skills

Proven ability to drive cultural and process change that fosters a collaborative approach between development and operations teams, Proven background in implementing and managing centralised logging solutions or similar platforms (e.g., Splunk, DataDog), Familiarity with distributed tracing tools (e.g., Jaeger, Zipkin) and Application Performance Monitoring (APM) solutions

Technologies

Terraform, Pulumi, CI/CD pipelines, Infrastructure-as-code (IaC), Kubernetes (K8s), Docker, Cloud providers (AWS, GCP, Azure), Monitoring, logging, distributed tracing, Observability, Go, Python, Typescript, Rust

Responsibilities

Build and maintain platform infrastructure using declarative IaC tools, Proactively manage the capacity of the infrastructure to consistently meet or exceed Service Level Objectives for latency, error rates, and availability, Act as first-line responders for critical system incidents, Triage, diagnose, and resolve complex production issues rapidly, Drive a culture of blameless post-mortems, Implement, maintain, and evolve the fully automated CI and CD pipelines, Implement and manage robust systems for monitoring (metrics), logging (centralised log aggregation), and distributed tracing to provide deep insights into application and infrastructure health

Seniority

Senior

Domain

Security, Infrastructure, Performance

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗