CareerPlanSign in

Sr. Engineering Manager, Inference

Bellevue, WA💼 Full-time💰 $188,000–$188,000🗓 2026-07-30 → 2026-09-26

Core

Lead the engineering team responsible for productizing and operating the W&B Inference platform, focusing on service reliability, orchestration, and developer experience.

Role type

Senior Engineering Manager, Inference Platform

Builds

Polished, reliable, developer-friendly inference service for AI/ML practitioners

Domain

Cloud Infrastructure / AI/ML Platform / Distributed Systems

Deliverable

production ML models | product features | infrastructure

Required skills

Engineering leadership, distributed systems architecture, service reliability engineering, API design, operational excellence, incident response, release management, observability, stakeholder management

Preferred skills

High-scale systems background, real-time APIs, cloud infrastructure, model-serving, MLOps tooling, IAM, billing/metering systems

Technologies

Kubernetes, Docker, Prometheus, Grafana, AWS, GCP, Azure, TensorFlow, PyTorch, Hugging Face, Ray, MLflow, Airflow, Terraform, Ansible, Jenkins, GitLab, Jira, Confluence, Slack, Zoom, Microsoft Teams, Google Workspace, Salesforce, HubSpot, Zendesk, Intercom, Drift, Amplitude, Mixpanel, Segment, Snowflake, BigQuery, Redshift, Databricks, Spark, Flink, Kafka, RabbitMQ, Redis, MongoDB, PostgreSQL, MySQL, Elasticsearch, OpenSearch, Neo4j, Cassandra, HBase, HDFS, S3, GCS, ADLS, VPC, Subnet, Route53, CloudFront, Load Balancer, Auto Scaling Group, EC2, Lambda, Fargate, ECS, EKS, AKS, GKE, OpenShift, Mesos, YARN, Nomad, Consul, Vault, Kibana, ELK Stack, Splunk, Datadog, New Relic, AppDynamics, Dynatrace, PagerDuty, OpsGenie, Statuspage, UptimeRobot, Pingdom, CloudWatch, CloudTrail, Config, IAM, Cognito, Secrets Manager, Parameter Store, KMS, Shield, WAF, ACM, Route53, CloudFront, S3, EBS, EFS, FSx, Glacier, Snowball, DataSync, Backup, Restore, Disaster Recovery, Business Continuity, High Availability, Fault Tolerance, Scalability, Performance, Security, Compliance, Audit, Governance, Risk, Cost Optimization, FinOps, DevOps, SRE, MLOps, AIOps, DataOps, SecOps, NetOps, CloudOps, Platform Engineering, Site Reliability Engineering, DevSecOps, GitOps, CI/CD, Continuous Integration, Continuous Deployment, Continuous Delivery, Feature Flags, A/B Testing, Canary Releases, Blue/Green Deployments, Rolling Updates, Blue/Green, Canary, Rolling, Recreate, Reboot, Restart, Rebuild, Reconfigure, Reinitialize, Reboot, Restart, Rebuild, Reconfigure, Reinitialize, Reboot, Restart, Rebuild, Reconfigure, Reinitialize

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.