We use cookies to improve your experience and measure how our site is used. Learn more.
Incidents get resolved, but the fixes that would stop them recurring slip down the backlog — and a few weeks later, a version of the same failure comes back. Most teams live with that. On a platform as critical as Buildkite's, we've decided not to, which is why we're creating a role whose whole remit is to own how we respond to incidents and design them out for good.
Buildkite runs production builds for teams like OpenAI, Anthropic, Uber, Shopify, Airbnb, Canva and Pinterest — software delivery in the critical path for over a billion daily users. When something breaks here, it breaks loudly, for people who can't afford it to.
Until now, Incident response has been carried by whoever was closest, in the gaps between everything else they own. It's held up. But "held up" is no longer the bar for a platform this critical. We'll be asking you to set the standard for how the whole engineering org detects, responds to, learns from, and designs out the incidents that matter — end to end.
You'd own incident management across engineering — the definition, process, tooling, and adherence — and make it stick.
In the moment, you're in the room on the high-impact incidents, driving clarity, coordination and a clean resolution. Between incidents — where the real work lives — you're running postmortems that go somewhere, making sure the followups get done every time, and turning what each one taught us into the changes that make the next incident smaller.
You'd work across every product team plus Security and Support, holding one consistent resilience bar and bringing engineering managers with you. And you'd coach engineers on how to run an incident, communicate through it, and own the outcome.
You've built or materially lifted incident management capability somewhere before, and you've got a clear, earned point of view on what good looks like.
You can read the code, dig into the telemetry, and lead a real investigation — not just chair the call. You're calm when it's sustained and messy: you make incidents smaller, not louder. And you can hold a firm bar across teams that don't report to you without softening the standard or bruising the relationship — the diplomacy to find common ground and the spine to keep the line.
You'll be at home with observability tooling (Datadog, Honeycomb), AWS, Terraform, and the failure modes distributed systems throw up at scale. Our stack is Ruby on Rails, PostgreSQL, Kafka and Redis on AWS — you won't need all of it on day one.
The one thing we won't move on: you've owned incident response for production systems at real scale for a number of years, meaning you've experienced a breadth of different challenges and scenarios. It's this knowledge and real world experience we're hoping to benefit from.
Our engineering teams are based across ANZ and US-Pacific time zones — a conscious decision that lets us move quickly with real overlap and minimise fully async work. So while Buildkite is fully remote, we don't hire everywhere: if you're applying from outside these regions, we're not currently in a position to hire you, and can't offer sponsorship.
Every application gets a response. We're wired for velocity, so you won't be left waiting on us. If this is the problem you've been wanting to get your hands on, apply — or reach out with questions first.
At Buildkite, we value diversity and celebrate all types of skills, backgrounds, and experiences. We’re dedicated to fostering an inclusive environment and providing reasonable accommodations throughout our recruitment process.
If you need any accommodations or support during the application or interview process, please reach out to us at accommodations@buildkite.com.
Health Insurance
Wellness Programs
401(k) / Retirement Plan
Profit Sharing
Equity / RSUs
Home Office Stipend
Company Equipment
Coworking Allowance
Internet Stipend
Learning & Development Budget
Conference Attendance
Mentorship Program
Buildkite runs a six-stage process: an application, a recruiter screen, a role-specific Technical Interview (which may include a hands-on task), a Craft Interview assessing four competencies (Productivity, Team Player, Customer Obsessed, and Curiosity), a Values Interview, and a final outcome communicated by email. Because the company is fully remote, every stage happens over chat and video.
Buildkite is globally distributed and remote-first, so you can work from a home office, a co-working space, or anywhere you focus best. Current openings are concentrated in the ANZ region (Australia and New Zealand), the United States, and Canada.
Recent openings span engineering (platform, data, ML, and security roles up to Principal and VP level) alongside go-to-market and customer-facing positions such as Account Executive, GTM Engineer, Customer Success Manager, and Technical Account Manager.
Buildkite is a small but growing team of roughly 150 people, deliberately kept lean on the belief that small teams can achieve big things. It is venture-backed, having raised about $41M across a 2020 Series A and a 2022 Series B.
People who do well here are comfortable working async in a small, globally distributed team, enjoy owning a wide range of projects, and value flexible hours over fixed office time. The culture centers on transparency, quality, collaboration, and sustainable growth.
PTO / Vacation
Sick Days
Parental Leave (Primary)
Parental Leave (Secondary)
Flexible Schedule
Flexible Hours
Work From Anywhere