All jobs

Engineering Manager - Site Reliability *EU/UK remote* (m/f/d)

Pliant2h ago
RemoteRemoteFull-timeManager Level5+ yrs exp

Top focus

Engineering ManagerVp EngineeringSreSenior Engineering Manager

ABOUT US Pliant is a European fintech specializing in B2B payment solutions. Our modular, API-first platform helps businesses streamline spending, improve cash flow, and integrate payments into their financial workflows. Designed for industries with complex payment needs, such as travel and fleet, Pliant enables greater efficiency, control, and profitability.

We serve two primary customer segments: Companies looking to optimize operational processes through intuitive apps and APIs, gaining control, automation, and financial flexibility through extended credit lines. Businesses such as financial software platforms, ERP providers, and banks that want to launch or enhance their credit card offerings using Pliant's embedded finance and white-label solutions.

Founded in 2020 and headquartered in Berlin, Pliant supports over 4,000 businesses and more than 20 partners globally. As a licensed e-money institution (EMI), we issue credit cards in 11 currencies across more than 30 countries, helping companies streamline and simplify payments.

Learn more at www.getpliant.com ABOUT THE ROLE You're the first hire for a Site Reliability function that doesn't exist yet at Pliant. Today, reliability is a responsibility scattered across many different teams, each with their own priorities: on-call rotation is only just being introduced, there's no framework for SLOs, and nobody's job is reliability instead of firefighting it on the side of something else.

You'll standardize the practice at Pliant while hiring the engineers to run it, coaching engineering teams to shift reliability from a burden into an integral part of the software development lifecycle. WHAT YOU'LL DO Define the framework other teams use to set their own SLOs and error budgets, educating and supporting product teams along the way Own blameless post-mortems and root-cause fixes; repeat incidents are treated as a process gap, not a signal about whoever was paged Implement production readiness reviews so nothing new ships without one Close gaps in Datadog observability coverage, including missing alerts, dashboard blind spots, and noisy pages that erode trust in on-call Hire and build the team from the ground up, setting the technical and cultural bar for every engineer who joins after you WHAT YOU'LL BRING 7-10 years of engineering experience, including at least 3 years directly managing engineers A track record of hiring and developing engineers, with specific people you've levelled or promoted Hands-on production or reliability engineering background.

This is not a first management role, and you've carried a pager yourself Strong AWS and Terraform experience, comfortable working inside a managed infrastructure-as-code pipeline Experience building or running an on-call rotation and incident management process, not just participating in one Strong platform observability experience.

You know the difference between a good dashboard and a useless one Clear communication for a technical, cross-team audience A track record of pushing reliability practices upstream into product engineering teams, not just reacting to incidents after the fact Proficiency with AI-assisted development (Claude Code, Cursor).

Comfortable reviewing AI-written PRs as rigorously as any other THE FIRST YEAR The first few months are about building the on-call rotation and incident process from scratch, since neither exists yet, and hiring the first engineers onto the team By mid-year, you have SLOs defined for the services that matter most, a real incident review process, and at least one engineer on the team besides you By year one, Site Reliability is a function other teams actually route to, not something they route around, and repeat incidents are trending down because the root-cause fixes stuck STACK Terraform, Spacelift, AWS (including a dedicated PCI-scoped account), Datadog.

We're subject to PCI DSS, SOC 2, and ISO 27001, and Platform Core's migration toward Kubernetes will increasingly shape what reliability looks like here too. WHAT WE OFFER The opportunity to work in a growing team with big responsibilities that thrives on a strong exchange of knowledge and excellence Attractive remuneration Your choice of preferred OS, Windows or Mac Flat hierarchy and transparent communication in a relaxed, professional atmosphere Opportunity to develop your talent in a dynamic team with ambitious goals Flexibility and possibility to work remotely Pliant Card with monthly credit to explore the product and enjoy food with colleagues Pliant is an equal opportunity employer.

We welcome applications from people of all backgrounds, identities and abilities, and are committed to an inclusive hiring process. If you need any accommodations during the interview process, please let us know.

Required skills

AWSKubernetesTerraformDatadog
Posted on JobRush — the end-to-end AI job-search platform.