Principal Software Engineer, Site Reliability
Core
Design, engineer, and build SRE platform systems and capabilities with cutting-edge AI that other engineering teams depend on in their critical path.
Role type
Principal Software Engineer, Site Reliability
Builds
SRE platform systems, monitoring, alerting, cloud infrastructure, automated remediation, and AI-powered capabilities
Domain
Enterprise Software / Site Reliability Engineering / AI
Deliverable
production ML models | product features | infrastructure
Required skills
Architecting large-scale distributed commercial applications, building complex internal platforms adopted by 10+ teams, driving system adoption through direct integration work, building and maintaining complex AI-powered applications, proficiency in object-oriented languages (C#, C++, Java, Python), deep understanding of data structures and algorithms, multithreading and synchronization, asynchronous patterns, cloud programming, service-oriented and microservice architectures, HTTP and web services development, modern engineering practices (agile, CI/CD, DevOps), managing production Kubernetes infrastructure, working with cloud providers (Azure, AWS, GCP) and managed services, database backend experience
Preferred skills
Experience with Azure SQL, CosmosDB, Azure Data Lake, Power BI, MongoDB, MySQL, DynamoDB
Technologies
C#, C++, Java, Python, Kubernetes, Azure, AWS, GCP, AKS, GKE, Azure SQL, CosmosDB, Azure Data Lake, Power BI, MongoDB, MySQL, DynamoDB
Responsibilities
Design and build SRE platform systems with AI, participate in livesite monitoring rotations and handle escalations, drive availability and scalability improvements, onboard other teams onto platforms by writing integrations and removing friction, mentor Software Engineers, influence process improvements and best practices
Seniority
Principal, strategy & mentorship