Application Support Engineer
Core
Lead Site Reliability Engineering (SRE) initiatives to automate operational toils, reduce technical debt, and manage critical business systems across multi-cloud environments.
Role type
Senior Site Reliability Engineer (Team Lead)
Builds
Automated infrastructure, CI/CD pipelines, observability frameworks, and monitoring agents
Domain
Cloud Infrastructure, Site Reliability Engineering, DevOps
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Infrastructure as Code (IaC), Cloud automation, Python, PowerShell, Ansible, Terraform, Docker, CI/CD pipelines, Observability tools, Splunk, Technical debt reduction, Capacity planning
Preferred skills
Enterprise architecture, End-client technical discussions, Compliance management, Project management principles
Technologies
AWS, Azure, Docker, Jenkins, GitHub, Puppet, ARM templates, Splunk
Responsibilities
Lead SRE team to achieve operational goals, Design and implement observability frameworks, Automate operational tasks and reduce technical debt, Troubleshoot critical events and manage application availability, Generate reports and dashboards for performance analysis, Coordinate with stakeholders and external partners for support and development issues
Seniority
Senior, hands-on IC with leadership responsibilities