Senior Platform Engineer (f/m/d) @ A1 Competence Delivery Center
Core
Building and operating an enterprise AI-as-a-Service platform integrating DataOps, MLOps, and LLMOps on Kubernetes.
Role type
Senior Platform Engineer (Kubernetes & AI Infrastructure)
Builds
Containerized AIaaS, DataOps, MLOps, and LLMOps services on Kubernetes
Domain
Telecommunications / Cloud Infrastructure / AI Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Kubernetes operations, Helm, CI/CD, scripting (Python/Go/Bash), API integration, observability, troubleshooting
Preferred skills
Cilium, CRI-O, Traefik, cert-manager, Kyverno, Falco, Harbor, Velero, Prometheus, Grafana, Loki, OpenTelemetry, Airflow, Airbyte, Meltano, OpenMetadata, MLflow, Feast, JupyterLab, LangFlow, vLLM, RAG platforms, GPU scheduling, NVIDIA GPU Operator, sovereign-cloud providers, workload portability assessment
Technologies
Kubernetes, Helm, Python, Go, Bash, Prometheus, Grafana, Loki, OpenTelemetry, Exoscale
Responsibilities
Deploy and maintain containerized AIaaS, DataOps, MLOps, and LLMOps services on Kubernetes; Integrate services such as workflow orchestration, data integration, metadata management, model tracking, feature stores, notebooks, model serving, RAG, and LLM observability; Develop and maintain Helm charts, Kubernetes manifests, configuration templates, and deployment pipelines; Establish repeatable deployment patterns across DEV, test, and production environments; Configure namespaces, service accounts, RBAC, resource quotas, network policies, secrets, certificates, and application ingress; Integrate platform services with enterprise APIs, databases, object storage, identity providers, monitoring systems, and internal data sources; Implement application-level observability using metrics, logs, traces, health probes, dashboards, and actionable alerts; Support vulnerability remediation, container-image governance, certificate renewal, secrets rotation, backup integration, and platform hardening; Diagnose failures spanning Kubernetes workloads, service configuration, APIs, identity, networking, storage, and external dependencies; Produce deployment documentation, technical runbooks, troubleshooting guides, and handover material; Review vendor deliverables and identify hidden infrastructure assumptions, privileged dependencies, portability gaps, and operational risks; Help translate pilot implementations into repeatable, supportable production services
Seniority
Senior, hands-on IC
