Platform Engineer
Platform Engineer(Hybrid)
incident.io
As a Platform Engineer at incident, you own everything that isn't the product itself: infrastructure, CI/CD pipelines, databases, and the systems that let our product engineers ship fast without thinking twice about what's underneath them. It's a genuinely varied role. One day you're tracing a scaling issue through a database, the next you're rethinking a Terraform module, the next you're pairing with a product engineer to unblock something that's been quietly slowing them down.
About incident.io
incident.io is the software reliability platform trusted by engineering teams at Netflix, Etsy, and 2,000+ companies who can't afford downtime. We connect On-call, Investigations, and Incident Response in one platform, powered by Nexus, a living AI model of your production environment—your runbooks, your services, and every incident you've ever run.
About The Role
Being a Platform Engineer at incident means being the detective every other engineer wishes they could clone, tracing a mystery from a flaky load balancer to a bad Terraform diff, turning "why is this broken" into "here's exactly why, and here's how we stop it happening again."
You'll work closest with our product engineers, day in and day out. You're part of a small, three-person platform team, which means your fingerprints are on everything. There's no hiding in a big team here, and no waiting for someone else to pick things up.
Responsibilities
- You take real ownership. You don't wait to be told what to fix, you see the thing that's slowing everyone down and you go fix it.
- You build tooling like a product engineer would. You think about the person using it, not just whether it technically works.
- You're genuinely great at observability. You don't just set up dashboards because you're told to, you want to actually see what's happening in the system, and you get restless when you can't.
- You've got a knack for unpicking confounding issues, the kind where the bug could be in the app, the database, the load balancer, or a Terraform file, and you're the person who enjoys following the thread until you find it.
- You care, properly, about making other engineers' lives better. The measure of a good day for you isn't just "it works," it's "that was annoying for people, and now it isn't."
- You keep half an eye on what's new in security, cloud, and developer experience, not because you have to, but because you'd be reading about it anyway.
- You're comfortable with high autonomy. Nobody's going to hand you a detailed spec, you're trusted to figure out what needs doing and just do it.
Benefits
- Private medical insurance. Seriously good cover - we want you and the people you love to be looked after.
- Competitive annual leave. Showing up at your best requires switching off, and we make sure you have time to do that.
- First Friday of every month off. Yes, seriously.
- Enhanced pension. We put real money in, because future-you deserves better than an afterthought.
- Meaningful equity. We're rapidly scaling, and everyone who helps shape the outcome should share in it.
- Unlimited AI spend. For everyone, not just engineers. We're all-in on AI across the company, and we expect you to be too.
- Generous parental leave. The early days with a new baby matter more than anything we're doing here, and we want you to be present for them.
- Two budgets that have your back. £1000 to invest in your setup, £500 a year to invest in yourself.
More Roles
Other roles you might like
Platform Engineer
Platform Engineer(On- Site)
Method Resourcing
This is a hands-on engineering role suited to someone who enjoys building and improving cloud platforms, automating infrastructure and driving engineering best practices. You'll play a key role in shaping platform standards, improving operational resilience and enabling the continued growth of a business-critical technology platform.
Site Reliability Engineer
Site Reliability Engineer
Moniepoint Group
We are seeking an experienced SRE to engineer the reliability of our highly distributed platform. You will combine deep knowledge of distributed systems with strong coding skills to define SLOs, lead incident response, and build automation and self-healing mechanisms into our systems.