Software Engineer (Backend, Python) - Content Understanding
Core
Design, build, and optimize distributed systems to extract, enrich, and process metadata from millions of documents, images, and audio content using machine learning and LLMs.
Role type
Software Engineer II (Backend, Python)
Builds
Scalable metadata pipelines and infrastructure for content discovery across Scribd, Slideshare, Everand, and Fable.
Domain
Digital content, machine learning, distributed systems, cloud infrastructure.
Deliverable
production ML models | infrastructure
Required skills
Python, Scala, Ruby, distributed systems design, AWS (ECS, EKS, Lambda), Terraform, Spark/Databricks, system optimization, automated validation
Preferred skills
LLM integration, ML model deployment, public cloud experience (Azure/GCP)
Technologies
Python, Scala, Ruby on Rails, Airflow, Databricks, Spark, AWS, Terraform
Responsibilities
Design and build scalable systems for metadata extraction and enrichment; Integrate LLMs for summarization, classification, and extraction; Optimize and refactor existing systems for performance; Ensure data accuracy through automated validation; Participate in code reviews; Manage and maintain data pipelines and security infrastructure
Seniority
Mid-level, hands-on IC