Staff Software Engineer, Model Infrastructure
Core
Lead the design and development of the systems that power every AI request at Harvey, ensuring high availability, low latency, and operational excellence for AI inference.
Role type
Staff Software Engineer (Model Infrastructure)
Builds
Model Infrastructure platform, Unified Model Controller (UMC), Provider Platform, Observability & Cost Platform, AI Platform Foundation
Domain
AI Infrastructure, Distributed Systems, Cloud Infrastructure
Deliverable
production ML models
Required skills
Go, Java, Python, Rust, C++, distributed systems, cloud infrastructure, networking, observability, technical leadership, cross-functional collaboration
Preferred skills
AI infrastructure, LLM serving, machine learning platforms, model routing, inference gateways, policy-based serving systems, Kubernetes, service mesh, large-scale observability, SRE best practices, data infrastructure (Kafka, Spark, Flink, Airflow, Iceberg), GPU infrastructure, model training platforms
Technologies
Go, Java, Python, Rust, C++, Kubernetes, Kafka, Spark, Flink, Airflow, Iceberg, OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten
Responsibilities
Lead design and implementation of the Model Infrastructure platform; Build systems for high availability, low latency, and operational excellence; Design and improve the Unified Model Controller (UMC) and Model Selector; Develop systems for model provisioning, capacity management, and failover; Integrate new model providers and maintain APIs/SDKs; Improve observability through dashboards, alerting, and telemetry; Partner with Product Engineering on model launches and monitoring; Drive infrastructure efficiency through capacity planning and cost optimization; Collaborate with AI Research on evaluation and deployment foundations; Lead cross-functional initiatives and mentor engineers.
Seniority
Staff, hands-on IC with leadership