Skip to main content
Back to jobs

Senior Production Engineer

Software Engineer - Front End - Studio (Core) - Greenhouse

London 9/13/2026 Full-time
Apply on source ↗ Save

&lt;p&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;&lt;strong&gt;About Clear Street:&lt;br&gt;&lt;/strong&gt;&lt;/span&gt;&lt;/p&gt; &lt;div&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;Clear Street’s mission is to give every sophisticated investor access to every asset, in every market, through a unified platform built for speed, transparency and scale.&lt;/span&gt;&lt;/div&gt; &lt;div&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.&lt;/span&gt;&lt;br&gt;&lt;br&gt;&lt;/div&gt; &lt;div&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;For more information, visit&amp;nbsp;&lt;a href=&quot;https://clearstreet.io/&quot; target=&quot;_blank&quot; data-saferedirecturl=&quot;https://www.google.com/url?q=https://clearstreet.io&amp;amp;source=gmail&amp;amp;ust=1766166763873000&amp;amp;usg=AOvVaw3GfALVHjsW8XaQSab-3CbX&quot;&gt;https://clearstreet.io&lt;/a&gt;.&lt;/span&gt;&lt;/div&gt; &lt;p&gt;&lt;strong&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;The Role&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;As a Production Engineer, you sit at the intersection of software reliability and operational &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;excellence. You own the health, resilience, and recovery of our production systems—while &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;spending equal energy innovating solutions that eliminate human toil, reduce incident blast &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;radius, and raise the reliability bar across the entire platform.&amp;nbsp; &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;You will partner closely with engineering, operations, and business teams to understand daily &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;pain points and translate them into lasting automated solutions. Half your time is spent in the &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;trenches—supporting production, responding to incidents, and deeply understanding how our&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;systems behave under real conditions. The other half is yours to build: automation, tooling, and &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;observability platforms that make tomorrow&amp;amp;#39;s on-call shift meaningfully easier than today&#39;s.&amp;nbsp;&amp;nbsp;&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;strong&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;You will work on challenges like:&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Design and build comprehensive monitoring and observability platforms that surface the &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;right signal at the right time—eliminating alert fatigue and accelerating root-cause &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;analysis.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Develop intelligent automation and self-healing capabilities that diagnose issues, trigger &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;recovery workflows, and reduce mean time to recovery (MTTR) without manual&amp;nbsp;&lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;intervention.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Analyze incidents, identify systemic trends, and engineer solutions that prevent entire &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;classes of failures from recurring.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Build reusable runbooks, diagnostic tooling, and recovery playbooks that turn tribal &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;knowledge into scalable platform capabilities.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Create golden-path operational workflows—making the safest, most reliable path also &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;the easiest one for engineering teams to follow.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Partner with Platform Engineering to influence CI/CD pipelines, deployment safety, and &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;infrastructure resilience from a production reliability perspective.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Champion Infrastructure as Code, GitOps, and SRE best practices while helping teams &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;adopt modern engineering workflows.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Continuously measure production health through SLIs/SLOs/SLAs, and drive &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;engineering priorities based on reliability data.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Explore emerging technologies—including AI-assisted diagnostics and developer &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;tooling—that transform how we operate production systems.&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;The Team&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;We believe resilient systems are built by engineers who understand them end to end.&amp;nbsp; &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;Our Production Engineering team is the first and last line of defense for our production platform.&amp;nbsp; &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;We treat reliability as a product, with uptime and engineer experience as our north stars. We &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;combine the discipline of SRE with a builder&#39;s mindset: when we see a recurring problem, we &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;build a solution—not a workaround.&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;You will work across every engineering and operations team to understand failure modes, quantify &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;reliability gaps, and build platform capabilities that scale with the organization. Whether it&#39;s &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;reducing MTTR from hours to minutes, building self-service diagnostic tools, or designing &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;proactive alerting that catches issues before customers notice, your work will have immediate, &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;measurable impact.&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;If you&#39;re passionate about making production systems invisible to end users—and you get &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;energy from both firefighting and building the systems that make fires less likely—you&#39;ll thrive &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;here.&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;What We&#39;re Looking For&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;We&#39;re looking for engineers who combine operational instinct with a builder&#39;s discipline.&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;You should have:&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Strong hands-on Python skills—this is your primary language for automation and tooling.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Experience in SRE, Production Engineering, Platform Engineering, or a related discipline &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;with direct production ownership.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Proven track record of building automation and diagnostic tooling that improved recovery &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;times or reduced operational toil.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Deep familiarity with cloud-native technologies—Kubernetes, containers, distributed &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;systems—and how they fail in production.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Experience with observability platforms such as Datadog, and a strong intuition for what &quot;&lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;good&quot; monitoring looks like.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Exposure to Infrastructure as Code (Terraform) and GitOps-based deployment &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;workflows (ArgoCD, GitHub Actions, or similar).&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Familiarity with the broader technology stack: Java, Go, Kafka, Redis, Snowflake, and &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;Postgres.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Strong analytical and problem-solving skills—you thrive on ambiguous, high-stakes &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;production problems.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● A product mindset applied to operational tooling: you think about usability, adoption, and &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;documentation when building internal solutions.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Excellent communication skills and the ability to work fluidly across engineering, &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;operations, and business stakeholders.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Self-starter mentality—you identify opportunities, take initiative, and deliver with minimal &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;supervision.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Curiosity and a continuous learning mindset; fintech or financial industry background is a &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;plus.&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;The Technology You&#39;ll Work With&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;You&#39;ll operate and build on a modern cloud-native platform that includes:&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Kubernetes &amp;amp; AWS&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Terraform &amp;amp; ArgoCD&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● GitHub Actions&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Kafka, Redis&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● PostgreSQL &amp;amp; Snowflake&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Datadog&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Python, Go, Java&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● gRPC &amp;amp; Protobuf&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Internal Platform APIs and Developer Tooling&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;strong&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;What Success Looks Like&lt;/span&gt;&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;Within your first year, you&#39;ll have made a measurable impact on production reliability. Success l&lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;ooks like:&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Reducing mean time to detection (MTTD) and mean time to recovery (MTTR) across key &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;production systems.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Building automation that handles a meaningful percentage of incident scenarios without &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;human intervention.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Becoming a trusted subject matter expert for core platform components and their failure &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;modes.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Delivering observability and diagnostic tools that other engineers actually use and &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;depend on.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Establishing SLO baselines and driving engineering investment based on reliability data.&lt;/span&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;● Spending less of your time—and your teammates time—on repetitive manual toil.&lt;/span&gt;&lt;/p&gt; &lt;p&gt;&lt;br&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;Your impact won&#39;t be measured by the number of incidents you respond to—it will be measured &lt;/span&gt;&lt;span style=&quot;font-size: 12pt;&quot;&gt;by how reliably our systems run and how quickly we recover when they don&#39;t.&lt;/span&gt;&lt;/p&gt;<p>Find <a href="https://www.arbeitnow.co.uk">Jobs in United Kingdom</a> on Arbeitnow</a>

Infrastructure

Related on GOH