CareerPlanGet AI match score →

Forward Deployed Engineer - AI/ML Platforms

San Francisco💼 Full-time🗓 2026-07-02 → 2026-07-31

Core

Partner with enterprise AI organizations to design, deploy, and operate production-grade distributed AI workloads on Ray and cloud infrastructure.

Role type

Forward Deployed Engineer (AI/ML Platforms)

Builds

Scalable AI platforms and distributed AI applications for enterprise customers

Domain

Cloud Infrastructure, Distributed Systems, Machine Learning Operations

Deliverable

production ML models | infrastructure

Required skills

Kubernetes, Cloud Infrastructure (AWS/Azure/GCP), Infrastructure as Code, Python/Go/Java, Distributed Systems, Customer Engagement, Technical Architecture, Automation/Tooling

Preferred skills

Ray, Spark, Dask, Enterprise Consulting Experience

Technologies

Kubernetes, Terraform, Helm, GitOps, AWS, Azure, GCP, Ray, Spark, Dask

Responsibilities

Design and implement production-grade AI platform architectures on Kubernetes and public cloud infrastructure; Partner with customer teams to deploy, operate, and optimize distributed AI workloads; Lead implementation engagements covering installation, networking, security, and scaling; Troubleshoot complex distributed systems issues; Develop automation, tooling, and infrastructure-as-code; Collaborate with Product and Engineering to shape future platform capabilities; Share best practices through documentation and workshops.

Seniority

Mid-Senior, hands-on IC with customer-facing responsibilities

Rewrite
## About the Role At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. ## Responsibilities - Design and implement production-grade AI platform architectures on Kubernetes and public cloud infrastructure (AWS, Azure, and GCP). - Partner directly with customer platform, infrastructure, and ML engineering teams to deploy, operate, and optimize distributed AI workloads. - Lead implementation engagements that include platform installation, networking, security, observability, scaling, upgrades, and operational readiness. - Troubleshoot complex distributed systems issues spanning infrastructure, Kubernetes, networking, storage, and AI applications. - Develop automation, tooling, reference implementations, and infrastructure-as-code that accelerate customer success and improve repeatability. - Build trusted relationships with technical leaders, platform teams, and executive stakeholders, translating business objectives into robust technical solutions. - Collaborate closely with Product and Engineering to communicate customer requirements, identify product improvements, and shape future platform capabilities. - Share best practices through technical documentation, architecture guidance, workshops, and enablement. ## Requirements - 5+ years of experience in cloud infrastructure, platform engineering, DevOps, Site Reliability Engineering, or software engineering. - Experience building, deploying, or operating ML/AI platforms that support model training, inference, or large-scale data processing workloads. - Strong expertise with Kubernetes and containerized production environments. - Experience operating cloud infrastructure on AWS, Azure, or GCP, including networking, security, IAM, storage, and infrastructure automation. - Experience with Infrastructure as Code and modern DevOps tooling such as Terraform, Helm, GitOps, CI/CD pipelines, or similar technologies. - Strong software engineering skills in Python, Go, Java, or a comparable language, with experience building automation or production services. - Experience working directly with enterprise customers in consulting, professional services, field engineering, solutions architecture, or another customer-facing engineering role. - Excellent communication skills and the ability to work effectively with both executive and deeply technical stakeholders. ## Nice to Have - Familiarity with distributed computing frameworks such as Ray, Spark, Dask, or Kubernetes-native distributed systems. - A passion for solving difficult customer problems and building reusable technical solutions. - Willingness to travel as needed to work alongside strategic customers. ## Benefits - Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗