Senior ML Ops Engineer
Core
Design and maintain the AI Enablement Platform (AIP) to deploy, scale, and monitor machine learning models and AI agents for email security and insider risk management.
Role type
Senior individual contributor ML Ops Engineer
Builds
Self-service deployment tooling, resilient ML inference infrastructure, and developer platforms for ML practitioners.
Domain
Cybersecurity / AI Infrastructure
Deliverable
production ML models
Required skills
AWS SageMaker, Python, Java (Spring Boot), Terraform, Kubernetes, CI/CD pipeline design, autoscaling, observability (Grafana, Open Telemetry), ML lifecycle tools (MLflow), Triton, ONNX
Preferred skills
LLM serving infrastructure, agentic AI deployment patterns, FinOps, cybersecurity background, Model Context Protocol (MCP)
Technologies
AWS, SageMaker, EC2, ECS/EKS, SQS, S3, CloudWatch, IAM, Python, Java, Bash, Terraform, Docker, Kubernetes, Jenkins, GitHub Actions, Grafana, Open Telemetry, PyTorch, MLflow, Triton, ONNX
Responsibilities
Design config-driven deployment workflows; tune autoscaling and rate limiting for high-throughput inference; build observability stacks for model performance and drift; maintain CI/CD pipelines with automated testing; manage multi-region infrastructure via IaC; optimize cloud costs for ML workloads; support agent and LLM operations; mentor engineers on operational best practices.
Seniority
Senior, hands-on IC with architectural leadership