Principal Software Engineer- Site Reliability
Core
Design, engineer, and build SRE platform systems and capabilities with cutting-edge AI that other engineering teams depend on in their critical path.
Role type
Principal Software Engineer (Site Reliability)
Builds
SRE platforms, observability tools, and cloud access policy systems
Domain
Enterprise Software / Cloud Infrastructure / AI
Deliverable
production ML models | infrastructure
Required skills
Architecting large-scale distributed commercial applications, building complex internal platforms adopted by 10+ teams, driving system adoption through direct integration work, building and maintaining complex AI-powered applications, proficiency in object-oriented languages (C#, C++, Go, Python), deep understanding of data structures and algorithms, multithreading, synchronization, asynchronous patterns, cloud programming, service-oriented and microservice architectures, HTTP applications, web services development, modern engineering practices (agile, CI/CD, DevOps), managing production Kubernetes infrastructure, experience with cloud providers (Azure, AWS, GCP) and managed services (AKS, GKE), database backend experience (SQL, NoSQL, Data Lakes)
Preferred skills
None explicitly stated as preferred (all listed skills are framed as requirements or 'musts')
Technologies
C#, C++, Go, Python, Kubernetes, Azure, AWS, GCP, AKS, GKE, Azure SQL, CosmosDB, Azure Data Lake, Power BI, MongoDB, MySQL, DynamoDB
Responsibilities
Design and build SRE platform systems with AI, participate in livesite monitoring rotations and handle escalations, drive availability and scalability improvements, ensure technical deliverables meet reliability and performance expectations, onboard other teams by writing integrations and removing friction, ship early and iterate based on feedback, drive task planning and staffing, mentor Software Engineers, influence process improvements
Seniority
Principal, hands-on IC with strategic impact
