Machine Learning Platform Engineer
Bjak
<p style="min-height:1.5em"><strong>About A1</strong></p><p style="min-height:1.5em">There are over 5 billion users using basic applications today such email, notes, tasks that are not AI-native. Our mission is to build a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organising and workflows, with minimal prompting.</p><p style="min-height:1.5em">Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. The system must handle multi-step reasoning, interact with external tools, and remain reliable despite non-deterministic model behavior. Our objective is to help users complete tasks daily enjoyable with over ~90%* reduced time.</p><p style="min-height:1.5em"><strong>About the Role</strong></p><p style="min-height:1.5em">As an ML Platform Engineer, you will build the infrastructure and systems that power A1's AI capabilities.</p><p style="min-height:1.5em">You will design and operate the systems behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement.</p><p style="min-height:1.5em">You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence.</p><p style="min-height:1.5em"><strong>Focus</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Build and operate the ML infrastructure and platforms powering A1’s AI products</p></li><li><p style="min-height:1.5em">Design systems for model training, evaluation, deployment, inference, and experimentation</p></li><li><p style="min-height:1.5em">Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads</p></li><li><p style="min-height:1.5em">Improve reliability, scalability, latency, and cost efficiency of AI systems</p></li><li><p style="min-height:1.5em">Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement</p></li><li><p style="min-height:1.5em">Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster</p></li><li><p style="min-height:1.5em">Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions</p></li><li><p style="min-height:1.5em">Build production observability, monitoring, tracing, and alerting for AI/ML workloads</p></li><li><p style="min-height:1.5em">Improve AI systems across reliability, scalability, latency, throughput, and cost</p></li><li><p style="min-height:1.5em">Identify bottlenecks across the ML stack and continuously improve system performance</p></li><li><p style="min-height:1.5em">Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure</p></li></ul><p style="min-height:1.5em"><strong>Tech Stack</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Python</p></li><li><p style="min-height:1.5em">PyTorch / JAX</p></li><li><p style="min-height:1.5em">LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM</p></li><li><p style="min-height:1.5em">Cloud infrastructure</p></li><li><p style="min-height:1.5em">Distributed systems</p></li><li><p style="min-height:1.5em">ML/data pipelines and workflow orchestration</p></li><li><p style="min-height:1.5em">GPU infrastructure and performance tooling</p></li><li><p style="min-height:1.5em">Vector databases and retrieval infrastructure</p></li></ul><p style="min-height:1.5em"><strong>Ideal Experience</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Strong software engineering fundamentals and experience building production systems</p></li><li><p style="min-height:1.5em">Experience building ML infrastructure, platforms, or production machine learning systems</p></li><li><p style="min-height:1.5em">Experience with model deployment, inference, evaluation, or data pipelines</p></li><li><p style="min-height:1.5em">Strong understanding of distributed systems and system reliability</p></li><li><p style="min-height:1.5em">Ability to write clean, maintainable, production-quality code</p></li><li><p style="min-height:1.5em">Comfortable working in ambiguous, fast-moving environments</p></li><li><p style="min-height:1.5em">Bias toward ownership, experimentation, and continuous improvement</p></li></ul><p style="min-height:1.5em"><strong>Outcomes</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">AI infrastructure reliably supports production workloads at scale</p></li><li><p style="min-height:1.5em">Models can be trained, evaluated, deployed, and improved efficiently</p></li><li><p style="min-height:1.5em">Inference systems deliver strong latency, throughput, reliability, and cost efficiency</p></li><li><p style="min-height:1.5em">ML pipelines are reproducible, observable, maintainable, and robust</p></li><li><p style="min-height:1.5em">Model and infrastructure regressions are detected quickly and diagnosed efficiently</p></li><li><p style="min-height:1.5em">Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt for every AI product</p></li><li><p style="min-height:1.5em">The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge</p></li></ul><p>Find more <a href="https://www.arbeitnow.co.uk/english-speaking-jobs">English Speaking Jobs in United Kingdom</a> on Arbeitnow</a>