Lead Infrastructure Engineer - Storage
Core
Lead infrastructure engineer responsible for operating, tuning, and automating enterprise storage platforms (block, file, object) using AI-assisted workflows to ensure resiliency, security, and auditability.
Role type
Senior IC infrastructure engineer (storage)
Builds
Production storage services, automation pipelines, and AI-driven observability tooling
Domain
Financial services infrastructure, enterprise storage systems
Deliverable
production ML models | infrastructure
Required skills
Storage fundamentals (RAID, erasure coding, replication, snapshots, tiering, IOPS, latency, multipathing, SAN/NAS, object storage), Linux system performance and kernel concepts, scripting (Python, Go, Bash), observability tooling (Prometheus, Grafana, ELK, Splunk, Datadog, OpenTelemetry), incident management, AI-assisted infrastructure analysis and validation
Preferred skills
Kubernetes storage and stateful workload operations, Infrastructure as Code (Terraform, CloudFormation, Ansible, Chef, Puppet), telemetry pipelines (Kafka), ITSM platforms (ServiceNow), backup/disaster recovery strategy (RPO/RTO), security controls (KMS, HSM, secrets management), self-service platform design
Technologies
AI, OpenSearch, Ansible, Bash, Cloud, Ceph, Datadog, ELK, Grafana, HSM, IBM, Kafka, Kubernetes, LLM, Linux, Network, OpenTelemetry, Prometheus, Puppet, Python, ServiceNow, Splunk, Terraform, CI/CD
Responsibilities
Own and improve SLOs, SLIs, error budgets, and operational excellence for storage services; lead incident response for storage outages and performance degradation; create and maintain runbooks, escalation paths, and standardized operating procedures; operate and enhance block, file, and object storage platforms across on-premises and cloud environments; perform performance tuning, capacity planning, lifecycle management, and resiliency testing; build automation for provisioning, patching, upgrades, replication, backup, and compliance checks; implement AI-driven observability and AIOps capabilities including telemetry correlation and LLM-assisted incident workflows