Senior MLOps Engineer – Biomedical Data London – Hybrid (3 days a week in the office)
As a Senior MLOps Engineer, you will play a critical role in ensuring AI models successfully transition from research into reliable production systems. Your work will underpin models that support key scientific and portfolio decisions, helping researchers determine which diseases to target, which patient populations may benefit most and where future therapies should be focused.
This is an opportunity to own the lifecycle of cutting-edge machine learning systems you'll be responsible for the deployment, monitoring and ongoing operation of production machine learning models, ensuring they remain scalable, reliable and reproducible throughout their lifecycle.
You will take ownership of production models, managing everything from experiment tracking and model registration through to deployment, monitoring, retraining and continuous improvement.
This role is ideal for someone who enjoys solving complex operational challenges, building robust ML infrastructure and enabling research teams to deliver AI at scale.
Key Responsibilities - Own the operational lifecycle of production machine learning models, from deployment through to monitoring, retraining and retirement.
- Ensure experiment tracking and model registry platforms are used consistently, maintaining complete provenance from training data through to production models.
- Configure, optimise and troubleshoot distributed training and fine-tuning workloads across large-scale compute environments.
- Work closely with ML Engineers during model handovers, reviewing technical documentation before accepting operational ownership.
- Deploy and manage model serving infrastructure, optimising configurations to deliver reliable performance for downstream users.
- Monitor production systems, identify issues proactively and implement improvements to maximise reliability and performance.
- Champion MLOps best practices.
Requirements - PhD in Machine Learning, Computer Science, Software Engineering or a related discipline – Plus 3-6 years of post-study work experience working with biomedical data including genomics, multimodal biological data or large-scale foundation model training, would be highly advantageous.
- Extensive experience deploying, operating and supporting production machine learning workloads.
- Experience working with distributed training technologies such as PyTorch Distributed, DeepSpeed, FSDP or Ray Train.
- Hands-on experience with experiment tracking and model registry platforms such as MLflow, Weights & Biases or similar technologies.
- Familiarity with CI/CD pipelines for machine learning, including tools such as GitHub Actions or cloud-native workflow platforms.
- A solid understanding of cloud infrastructure, including compute, networking and storage for large-scale ML workloads.
- Experience supporting foundation model training, with knowledge of scaling, parallelisation and memory optimisation.
- Familiarity with infrastructure-as-code tools such as Terraform or equivalent cloud-native solutions.
This is a great opportunity to work alongside a talented team of AI scientists and ML engineers to support break throughs in drug discovery. Apply today to be considered.