Staff Production Operations Engineer
Core
Lead application operations, production support, and system architecture for public-facing services in AWS and container environments to ensure availability, resiliency, and scalability.
Role type
Staff individual contributor system architect and production operations engineer
Builds
Reliable production services, automation frameworks, CI/CD gates, and observability tools for millions of users
Domain
Cloud infrastructure, container orchestration, and distributed service operations
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Distributed service architecture, Unix/Linux internals, networking (TCP/IP, HTTP/HTTPS, DNS), AWS service operations, Kubernetes/EKS/UKS, Python/Go/Java, Infrastructure as Code (Terraform), CI/CD, observability (Datadog, CloudWatch, Splunk), incident management, SQL/NoSQL
Preferred skills
AI-assisted development workflows (Codex, Claude, Cursor), payment/commerce systems experience
Technologies
AWS (ALB, Route 53, API Gateway, Lambda, RDS, DynamoDB, ElastiCache), Docker, Kubernetes, EKS, UKS, Fargate, Python, Go, Java, GitHub, Ansible, Chef, Terraform, CloudFormation, Jenkins, Spinnaker, Datadog, CloudWatch, Splunk, Grafana, BigPanda, Slack, JIRA, ServiceNow, SQL, MySQL, Oracle, Snowflake
Responsibilities
Lead application operations and production support for internal and public-facing services; Define and evolve production operations architecture and readiness standards; Architect and scale tools, scripts, and automation frameworks; Establish CI/CD operational gates and governance mechanisms; Advance observability and event correlation; Partner with SRE and product teams to shift reliability left; Lead performance, capacity, infrastructure modernization, and cost optimization; Provide rotational on-call support and lead incident response and RCA.
Seniority
Staff, hands-on IC with broad technical ownership and mentorship