Site Reliability Engineer
Core
Ensures reliability, availability, and performance of web services and applications within a Python development squad by bridging development and operations.
Role type
Site Reliability Engineer (SRE)
Builds
Scalable, secure, and reliable systems using Python-based services and microservices
Domain
Financial Technology (CFD and spread betting broker)
Deliverable
production ML models | infrastructure
Required skills
Python (asyncio, FastAPI, Django, pytest, Poetry/pip), Cloud platforms (AWS, GCP, Azure), Container orchestration (Kubernetes, Docker), Infrastructure-as-Code (Terraform, CloudFormation), System design, Monitoring & Observability (Prometheus, OpenTelemetry, structlog), Incident management, Root cause analysis, CI/CD pipelines
Preferred skills
Profiling, benchmarking, dependency management
Technologies
Python, Prometheus, OpenTelemetry, structlog, Django, pytest, Poetry, pip, AWS, GCP, Azure, Kubernetes, Docker, Terraform, CloudFormation
Responsibilities
Design, implement, and maintain scalable systems; Build and manage monitoring, alerting, and logging systems; Develop and maintain automation tools and internal libraries; Partner with development squads to embed reliability; Conduct root cause analysis for incidents and implement corrective actions; Drive initiatives to improve system performance and scalability
Seniority
Mid-level, hands-on IC