Sustaining Engineer
Core
Deep technical investigations, root cause analysis, and long-term stability improvements for a large-scale distributed archive platform.
Role type
Senior Sustaining Engineer (Production Operations)
Builds
Archive platform (hybrid data center and AWS cloud infrastructure)
Domain
Cybersecurity / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
C#, Java, or C++ proficiency; distributed systems architecture; debugging complex production systems; full stack backend/API/database/infrastructure skills; CI/CD and automated testing; observability and logging; high availability design
Preferred skills
Microsoft Exchange and MAPI experience; AWS EC2, S3, RDS, IAM operations; messaging systems and search platforms; large scale storage systems; incident response leadership
Technologies
AWS (EC2, S3, RDS, IAM), C#, Java, C++, CI/CD tools, monitoring tools
Responsibilities
Lead deep technical investigations into production issues; Perform root cause analysis across application, database, infrastructure, and cloud layers; Implement code level and architectural fixes to improve system reliability; Partner with Support and Product teams on high severity customer escalations; Drive systemic improvements that reduce recurring incidents and operational toil; Improve observability, logging, and debugging capabilities; Participate in on-call rotation and respond to critical production incidents
Seniority
Senior, hands-on IC