Skip to main content
Back to jobs

Machine Learning Performance Engineer

Janestreet

London 9/22/2026 Full-time
Apply on source ↗ Save

&lt;p&gt;We are looking for an engineer with experience in low-level systems programming and optimisation to join our growing ML team.&amp;nbsp;&lt;/p&gt; &lt;p&gt;&lt;a href=&quot;https://www.janestreet.com/join-jane-street/machine-learning/&quot;&gt;Machine learning&lt;/a&gt; is a critical pillar of Jane Street&#39;s global business. Our ever-evolving trading environment serves as a unique, rapid-feedback platform for ML experimentation, allowing us to incorporate new ideas with relatively little friction.&lt;/p&gt; &lt;p&gt;Your part here is optimising the performance of our models – both training and inference. We care about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want to ensure our platform makes sense even at the lowest level – is all that throughput actually goodput? Does loading that vector from the L2 cache really take that long?&lt;/p&gt; &lt;p&gt;If you’ve never thought about a career in finance, you’re in good company. Many of us were in the same position before working here. If you have a curious mind and a passion for solving interesting problems, we have a feeling you’ll fit right in.&amp;nbsp;&lt;/p&gt; &lt;p&gt;There’s no fixed set of skills, but here are some of the things we’re looking for:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;An understanding of modern ML techniques and toolsets&lt;/li&gt; &lt;li&gt;The experience and systems knowledge required to debug a training run’s performance end to end&lt;/li&gt; &lt;li&gt;Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy&lt;/li&gt; &lt;li&gt;Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Computesight-systems and nsight-compute&lt;/li&gt; &lt;li&gt;Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS&lt;/li&gt; &lt;li&gt;Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization and asynchronous memory loads&lt;/li&gt; &lt;li&gt;Background in Infiniband, RoCE, GPUDirect, PXN, rail optimisation and NVLink, and how to use these networking technologies to link up GPU clusters&lt;/li&gt; &lt;li&gt;An understanding of the collective algorithms supporting distributed GPU training in NCCL or MPI&lt;/li&gt; &lt;li&gt;An inventive approach and the willingness to ask hard questions about whether we&#39;re taking the right approaches and using the right tools&lt;/li&gt; &lt;li&gt;Fluency in English&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;&amp;nbsp;&lt;/p&gt; &lt;p&gt;&lt;em&gt;If you&#39;re a recruiting agency and want to partner with us, please reach out to&amp;nbsp;&lt;/em&gt;&lt;a ----------<p>Find more <a href="https://www.arbeitnow.co.uk/english-speaking-jobs">English Speaking Jobs in United Kingdom</a> on Arbeitnow</a>

Machine Learning

Related on GOH