2 Devops Engineer Live Video
Core
Own and operate production infrastructure for AI agents automating patient access workflows in healthcare.
Role type
DevOps Engineer
Builds
Production infrastructure, GKE clusters, CI/CD pipelines, monitoring stacks
Domain
Healthcare automation, Cloud Infrastructure (GCP)
Deliverable
infrastructure
Required skills
Kubernetes, GitOps (Argo CD), Terraform, Helm, Prometheus, Grafana, Loki, Linux, Shell scripting
Preferred skills
Secrets management (Vault, Sealed Secrets), GCP/GKE expertise, Open source contribution
Technologies
GKE, Argo CD, Terraform, Helm, Prometheus, Grafana, Loki, HashiCorp Vault
Responsibilities
Operate production infrastructure with HA and autoscaling, Manage GitOps workflows, Maintain monitoring and alerting stacks, Implement infrastructure as code, Lead incident response and disaster recovery
Seniority
Mid-level, hands-on IC
Rewrite
## About the role
You'll be joining a specialised team at 100ms focused on healthcare automation using AI.
## About the company
100ms is building AI agents that automate complex patient access workflows in U.S. healthcare — starting with benefits verification, prior authorization, and referral intake in specialty pharmacy.
We help care teams reduce delays and administrative burden so that patients can start treatment faster. Our automation platform combines deep healthcare knowledge with LLM-based agents and robust ops infrastructure.
## What we offer
You'll be part of a small team at a fast-growing engineering-first startup.
You'll work with engineers across the globe with experience in video at places like Facebook and Hotstar.
You can grow as an individual contributor or as a team leader - freedom to set your own goals.
## Responsibilities
- Own and operate production infrastructure, multiple GKE clusters with HA, autoscaling, and observability.
- Manage GitOps workflows using Argo CD for automated, version-controlled deployments.
- Maintain and optimise monitoring & alerting stacks using Prometheus, Grafana, and Loki.
- Implement infrastructure as code using Terraform for GCP resources and Kubernetes manifests.
- Lead or support incident response, cluster upgrades, and disaster recovery procedures.
## Requirements
- Computer Science/Engineering or equivalent practical experience
- Minimum 3 years of hands-on experience with Kubernetes in a production environment.
- Strong knowledge of CI/CD pipelines and GitOps workflows using Argo CD or similar.
- Proficient in infrastructure automation using Terraform and Helm.
- Experience managing monitoring/logging stacks (Prometheus, Loki, Grafana, Alertmanager).
- Comfortable with Linux systems, shell scripting, and basic networking.
## Nice to have
- Prior experience with handling large infrastructure
- Knowledge of secrets management tools (e.g., HashiCorp Vault, Sealed Secrets).
- Prior experience with handling GCP and GKE
- Experience with open source contribution
- Ability to speak and write in English fluently and idiomatically
- Strong inclination to keep up-to-date with latest trends, learn new concepts, or contribute to open-source projects and would be eager to talk about ideas in internal or external forums
## Additional Information
At 100ms, we place a strong emphasis on in-office presence to promote collaboration and strengthen company culture.
Under the current policy, employees are expected to work from the office at least three days a week—Tuesday, Wednesday, and Friday—as an essential part of their role.
Website
https://www.100ms.ai/
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.