CareerPlanGet AI match score →

Infrastructure Engineer Devops Sde2

💼 Full-time🗓 2026-07-28

Core

Design, build, and operate scalable cloud infrastructure for a high-growth storytelling platform, including AI workloads, media delivery, and self-managed data layers.

Role type

Senior Infrastructure Engineer (DevOps/SRE)

Builds

Multi-region AWS infrastructure, CI/CD pipelines, media delivery systems, and operational substrate for AI/LLM workloads.

Domain

Cloud Infrastructure, AI/ML Operations, Media Streaming

Deliverable

production ML models | infrastructure

Required skills

AWS (ECS, VPC, IAM, multi-AZ), Terraform, Jenkins, Python, Ansible, Linux internals, TCP/IP, DNS, TLS, database operations (MySQL, MongoDB, Redis/Valkey), Kafka, ScyllaDB, GPU capacity planning, vLLM/TGI, observability (Prometheus, Grafana, OpenTelemetry), SLO management, incident response.

Preferred skills

Cloud security (ISO 27001, DPDP), cost optimization, media transcoding, LLM inference optimization, Langfuse-style tracing.

Technologies

AWS, ECS, EC2, VPC, IAM, Terraform, Jenkins, Ansible, Python, Cloudflare, Prometheus, Grafana, InfluxDB, RDS MySQL, MongoDB, Valkey, MSK, ScyllaDB Cloud, vLLM, TGI, OpenTelemetry, Langfuse

Responsibilities

Own AWS infrastructure via Terraform (ECS, VPC, IAM, autoscaling); build fast, safe Jenkins CI/CD pipelines; drive cost optimization; operate media delivery at scale (CDN, transcoding); manage self-managed data layer (MySQL, MongoDB, Valkey, MSK, ScyllaDB); stand up AI operational substrate (GPU, inference gateways, rate limits); partner with AI/ML teams on model self-hosting; ensure reliability (SLOs, incident response, postmortems).

Seniority

Senior, hands-on IC

Rewrite
## About the role Pratilipi is building a generational company in storytelling, and infrastructure is core to that bet. We're looking for an Infrastructure Engineer — a hands-on role embedded in a high-growth product team using AWS, CI/CD, our self-managed data layer, and the operational substrate for the AI workloads now landing in production. You'll help us self-host our own models and take the infrastructure multi-region as we expand globally. Expect a lot of breadth, with depth in a few things that matter. You won't be a specialist hiding behind a narrow remit — you'll move across AWS, CI/CD, databases, CDN, and AI infra, and go deep where it counts. ## What you'll do - Own our AWS infrastructure through Terraform — ECS on EC2, VPCs, IAM, autoscaling, rolling deploys. Clickops-free console as the target. - Build Jenkins pipelines that are fast, safe, and self-serve. Be in the deploy paths, not just the platform. - Drive cost optimisation. - Own media delivery at scale — Cloudflare CDN, image/video/audio transcoding, format/device delivery — tuned for latency, cache hit ratio, and egress cost. - Operate the self-managed data layer with service teams — RDS MySQL, MongoDB, Valkey, plus managed MSK and ScyllaDB Cloud. Lead migrations end-to-end (e.g. the ongoing Redis → Valkey). - Stand up the operational substrate for AI in production — GPU capacity, inference gateways, key rotation, rate limits, cost-per-request visibility. Build observability for LLM agents: traces, token/cost accounting, eval hooks, alerts on silent regressions. - Partner with the AI/ML team on self-hosting models — capacity planning, vLLM / TGI serving, canary rollouts, and cost/latency trade-offs vs managed providers. - Own reliability across services and infra — SLOs, alerting, incident response, blameless postmortems. Move teams from firefighting to proactive reliability. - Once you have the context, add the guardrails — IaC checks, deploy gates, paved paths — that make the right thing the easy thing and quietly remove whole classes of human error. ## What we're looking for - 4–6 years in DevOps, SRE, or infrastructure — production systems at meaningful scale, with ownership beyond tickets. - A problem solver with high agency. You reason from first principles, don't wait to be told, and dig in rather than deflect — whether it's a developer stuck on Terraform or an ML engineer asking for GPUs. - Strong AWS hands-on — ECS on EC2, VPCs, IAM, multi-AZ design — and proficiency with Terraform (you write modules others reuse). - Jenkins in production plus a real sense of developer experience in CI/CD — you've looked at deploy-time metrics and changed them. - Python, Ansible, and shell — you automate work rather than repeat it. - Operated databases in production — at least some of RDS MySQL, MongoDB, Redis/Valkey. Done migrations, failovers, and perf tuning, not just provisioning. Working familiarity with Kafka and Cassandra/ScyllaDB. - Strong grasp of Linux internals and networking fundamentals (TCP/IP, DNS, TLS, load balancing) and common failure modes. - Some hands-on exposure to LLM-based systems — inference endpoints, agent tracing, token/cost, or evals in CI. Not a researcher; you should reason clearly about latency, cost, and failure modes of LLM workloads. - Comfortable with cloud security fundamentals: IAM least-privilege, secrets management, network segmentation. Prior exposure to ISO 27001 / DPDP-style controls is a plus. ## Tech stack AWS · ECS (on EC2) · Terraform · Jenkins · Ansible · Python · Cloudflare · Prometheus · Grafana · InfluxDB · RDS MySQL · MongoDB · Valkey · MSK · ScyllaDB Cloud · growing: GPU inference · vLLM / TGI · OpenTelemetry GenAI · Langfuse-style tracing · multi-region AWS ## Security & Data Handling All employees are expected to handle sensitive data responsibly in compliance with the DPDP Act, ISO-27001:2022, and Pratilipi's internal security policies — ensuring data privacy, confidentiality, and NDA obligations at all times.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗