Rewrite
## Responsibilities
- Act as senior technical solution expert for complex issues spanning data pipelines, ML pipelines and/or AI applications, applying deep expertise in distributed systems.
- Analyse and troubleshoot production workloads at the code level, optimise for performance, reliability, latency, and cost.
- Diagnose and support Machine Learning and/or Large Language Model deployments, including real-time and batch inference, autoscaling, monitoring, logging, and alerting. Serve as a Subject Matter Expert guiding customers on experiment tracking, model registry, versioning, evaluation, labelling, tracing, and lifecycle observability.
- Provide high-quality support by guiding customers in leveraging Databricks AI to solve generative AI use cases & challenges, leveraging LLMs, MCP, AI Agents, RAG/Agentic RAG, APIs, vector embeddings, semantic search, Vector Search/Lakebase databases, context orchestration, memory management, and prompt engineering.
- Collaborate with internal teams to influence roadmap, product improvements and support business growth.
- Develop expertise in productionizing systems in Databricks and share your knowledge by contributing to wikis and other technical documentation, or by teaching our AI systems new skills, which will be used internally and externally by customers and partners.
## Requirements
- 8+ years of experience designing, building, and scaling Data, Machine Learning, and AI systems on-premises and in the cloud using Python, Scala, and Java in production environments, with expertise in Machine Learning and/or generative AI. Experience with cloud platforms (AWS, Azure, or GCP); familiarity with Databricks is a plus. Proficient in data engineering necessary for orchestrating end-to-end machine learning training pipelines, ideally with experience processing large datasets with Apache Spark.
- SME knowledge in feature engineering, ML frameworks, model training, model monitoring, drift detection, and retraining strategies. Proficient in working with algorithms and deep learning, along with NLP techniques.
- Prior experience building, designing or troubleshooting LLM-based Generative AI applications. Familiarity with agentic frameworks (e.g., LangChain, LangGraph etc). Expertise in context orchestration, including prompt design, memory management, retrieval systems, vector embeddings, semantic search, and tool integrations.
- Comprehensive Knowledge of MLOps and LLMOps with expertise in model evaluation, scoring, ranking, optimisation, training, validation, and packaging.
- Experience developing agent skills, plugins, and debugging with native AI capabilities is a plus.
- Prior support or customer-facing experience is not required for this role, but the ability and desire to develop excellent customer service skills are.
- Prior experience in Data Scientist, ML Engineer, or AI Engineer roles is highly valued.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field (or equivalent experience). Professional certifications are good to have.
## Nice to Have
- Experience with cloud platforms (AWS, Azure, or GCP);
- Familiarity with Databricks;
- Experience developing agent skills, plugins, and debugging with native AI capabilities;
- Professional certifications.
## Benefits
- Reporting to a TSE manager - you will be part of a world class global support engineering organization for Databricks, known for your technical depth and delivering impeccable customer service.
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.