about the company.
AI for Science.
about the team.
Information Technology.
about the role.
Our client is hiring a Post-Training / LLMOps Engineer to own the training infrastructure behind our next generation of scientific foundation models and to help drive the post-training that runs on top of it. This role is the systems backbone for post-training state-of-the-art open models — including leading frontier models such as Qwen, DeepSeek, GLM, Kimi, and MiniMax — and adapting them for scientific and drug-discovery use cases.
Your center of gravity is LLMOps and large-scale training systems: standing up and optimizing multi-node training across cloud and managed ML platforms, and making the difference between a job that "runs" and a job that runs fast, stably, and cost-efficiently at scale. Alongside that, you will work directly with post-training engineers and researchers across the full post-training stack (SFT, RFT/RL, continual pretraining) — setting up the environment, data, and parallelism strategy for fine-tuning runs, improving the systems that couple trainers, inference engines, and rollout workers, and wiring the resulting expert models into company's drug-discovery workflows, interfaces, and tools.
...
about the job.
- Design, provision, and operate large-scale distributed training environments for LLM post-training across multi-GPU, multi-node clusters.
- Stand up and optimize training on Kubernetes-based and managed ML platforms, including AWS SageMaker HyperPod / EKS, plain cloud Kubernetes (GKE, AKS, self-managed), Tencent TiOne, and similar platforms; abstract away platform differences so researchers can move between them with minimal friction.
- Configure and tune model and data parallelism strategies — data, tensor, pipeline, sequence/context, and expert (MoE) parallelism — to maximize throughput and fit models to available hardware.
- Increase training efficiency and minimize training cost through GPU utilization tuning, mixed-precision/quantization-aware training, communication/overlap optimization, optimal batch and sequence packing, spot/preemptible capacity strategies, and right-sizing of clusters.
- Own high-performance data and storage pipelines for training and benchmarking, including S3 / EFS / FSx for Lustre (and equivalents), data sharding/streaming, caching, and throughput tuning so data loading never bottlenecks the GPUs.
- Engineer robust checkpointing, fault tolerance, and recovery (including elastic and resumable training) so long multi-node runs survive node failures and preemptions.
- Build and maintain the environment and tooling around training jobs: container images, dependency/environment management, job schedulers, queueing, autoscaling, and reproducible run configuration.
- Build observability for training: GPU/throughput/cost dashboards, run telemetry, profiling, and alerting; diagnose stragglers, hangs, OOMs, optimization issues, and performance regressions.
- Establish MLOps/LLMOps best practices — CI/CD for training and serving, infrastructure-as-code, experiment tracking, artifact/model registries, and reproducibility standards.
- Design, run, and improve post-training workflows for LLMs, including SFT, RFT/RLVR/RLAIF, and continual pretraining for scientific and reasoning-heavy use cases.
- Build and optimize the infrastructure that couples trainers, inference engines, asynchronous rollout workers, and evaluation systems for RL-based post-training.
- Support the data-to-benchmark loop: help convert internal drug-discovery data into datasets usable for model training and prospective benchmarking, and operate the infrastructure for running benchmark/evaluation suites at scale.
- Analyze failures in fine-tuning runs, diagnose optimization or systems bottlenecks, and turn experimental findings into robust engineering improvements.
- Help integrate post-trained expert models into production drug-discovery workflows, serving stacks, and internal tools, in collaboration with platform and product teams.
skills and experience required.
- M.S., Ph.D., or equivalent relevant experience in Computer Science, Machine Learning, Distributed Systems, or a related quantitative/engineering discipline.
- 5+ years of hands-on experience spanning ML infrastructure / LLMOps and LLM training, or closely related areas.
- Deep, hands-on expertise with Kubernetes for large-model training, including at least one managed/large-scale platform such as AWS SageMaker HyperPod / EKS, GKE/AKS, Tencent TiOne, or comparable; comfortable across multiple platforms rather than locked to one.
- Strong practical understanding of distributed-training parallelism — data, tensor, pipeline, sequence/context, and expert parallelism — and how to combine them for throughput and memory efficiency.
- Hands-on experience with distributed-training frameworks and stacks such as PyTorch (DDP/FSDP), DeepSpeed/ZeRO, and the Hugging Face ecosystem; proven ability to debug and optimize multi-node runs rather than only launching baseline scripts.
- Fluency with cloud storage and high-throughput data systems for ML — S3, EFS, FSx for Lustre (and equivalents) — including data-loading optimization and storage/throughput tuning.
- Proven track record of increasing training efficiency and reducing training cost on real workloads (utilization, scaling efficiency, capacity strategy).
- Demonstrated experience setting up the environment, sequence, and data for post-training of open models, and operating the infrastructure for benchmarking post-trained model performance.
- Practical understanding of post-training methods — SFT and RL approaches such as GRPO, PPO, DPO, and reward modeling — and the systems implications of RL training loops (trainers, inference engines, rollout workers).
- Strong skills with containers and orchestration (Docker, Kubernetes) and infrastructure-as-code.
- Strong Python skills and solid engineering habits around reproducibility, automation, CI/CD, observability, and debugging at scale.