You make the difference between an AI system that ships once and one that ships every week without drama. You will own the MLOps and AI platform layer across engagements, treating reliability and observability as the product.
What you will do
- Build CI/CD and deployment paths for models, prompts, and AI services.
- Stand up monitoring for drift, cost, latency, and quality – with alerts that mean something.
- Own model and prompt versioning, rollback, and safe-deploy practices.
- Manage inference infrastructure cost and scaling across clouds.
- Pave the AI platform roads so forward-deployed engineers move faster.
What we are looking for
- 7+ years in platform, infrastructure, or MLOps roles.
- Strong with Kubernetes, Docker, Terraform/Pulumi, and at least one major cloud.
- Experience operating ML or LLM systems in production, including observability.
- Python plus the systems depth to debug an inference path under load.
- A reliability mindset and calm under incident pressure.
Nice to have
- Experience with LLM observability tooling and token-cost control.
- Background in regulated or high-availability environments.
Apply for this role
Tell us about yourself and attach your resume. We review every application and reply within a few business days.