We use cookies to improve your experience and measure how our site is used. Learn more.
Push a one-line fix. Then watch CI grind through forty minutes of tests, ninety-five percent of which never had a chance of touching what you changed. You already know the handful that mattered. The test suite doesn't — so it runs everything, every time, just in case.
That "just in case" is the most expensive habit in software delivery. Every engineering team pays it, because the alternative — knowing which tests actually matter for a given change — has been too hard to get right.
We're building the team that gets it right. This role sits at the centre of it.
Test Engine already ingests billions of test runs. We can see the tests, the code underneath them, and how the two move together — at a scale very few people ever get to work with. The raw material for the answer is already here. Nobody's turned it into predictions yet.
That's the step to take: for a given change, work out the slice of tests most likely to fail, and run only those. Get it right and teams stop re-running what hasn't changed, and spend that time where it counts — like fixing the two percent of tests most likely to break.
It's a genuinely difficult ML problem — sparse signal, cold-start on new repos, generalising across languages and frameworks, and latency tight enough to sit in the critical path. It's also close to a blank page. There's no ML org above you setting the direction — you'd set it. And not alone: we've just hired another ML engineer, so there's someone to think out loud with from day one.
Machine learning in Test Engine, end-to-end — the strategy, the architecture, and the models running in production.
That means shaping the whole path: pulling features out of code changes and test history, training and evaluating models, building the serving layer that keeps predictions fast, and closing the loop so the system keeps improving. You'd make the trade-offs that matter — accuracy versus latency, what happens when confidence is low — and build the platform underneath so the next model into production is quick and repeatable, not a one-off.
You've taken ML models the whole way — from rough idea to something running reliably in production, monitored and retrained, owned rather than handed off.
Two things matter more than any specific tool:
Day to day you'll live in Python and SQL, on AWS, with containerised workloads and data-at-scale tooling (Spark, Flink, or similar). Experience with code analysis, CI/CD systems, or ranking problems is a real head start — a bonus, not a bar.
The one thing we won't budge on: you've shipped and owned ML in production. Prototyped and handed off doesn't count here.
You're likely a strong fit if you:
This probably isn't the right role if you:
We'd rather you know that now than three interviews in.
Every application gets a response. If this is the problem you've been wanting to get your hands on, apply now, or reach out with questions first.
Our Engineering teams are based in the ANZ/PST region. This is a conscious decision, as it allows us to move quickly and minimise fully async work. So whilst Buildkite is a fully remote company, this doesn't mean that we hire in every location. Please be aware that if you're applying from outside of this region, we unfortunately aren't in a position to hire you. Currently Buildkite is not in a position to offer sponsorship.
At Buildkite, we value diversity and celebrate all types of skills, backgrounds, and experiences. We’re dedicated to fostering an inclusive environment and providing reasonable accommodations throughout our recruitment process.
If you need any accommodations or support during the application or interview process, please reach out to us at accommodations@buildkite.com.
Health Insurance
Wellness Programs
401(k) / Retirement Plan
Profit Sharing
Equity / RSUs
PTO / Vacation
Flexible Hours
Home Office Stipend
Company Equipment
Coworking Allowance
Parental Leave
Learning & Development Budget
Conference Attendance
Buildkite runs a six-stage process: an application, a recruiter screen, a role-specific Technical Interview (which may include a hands-on task), a Craft Interview assessing four competencies (Productivity, Team Player, Customer Obsessed, and Curiosity), a Values Interview, and a final outcome communicated by email. Because the company is fully remote, every stage happens over chat and video.
Buildkite is fully remote, so you can work from a home office, a co-working space, or wherever you focus best. Hiring is limited to two corridors: the ANZ region (Australia and New Zealand, described in some postings as APJ) and the US Pacific timezone. Buildkite states that applicants outside those regions cannot currently be hired, and that it is not in a position to offer sponsorship.
Recent openings span engineering (platform, data, ML, and security roles up to Principal and VP level) alongside go-to-market and customer-facing positions such as Account Executive, GTM Engineer, Customer Success Manager, and Technical Account Manager.
Buildkite is a small but growing team of roughly 150 people, deliberately kept lean on the belief that small teams can achieve big things. It is venture-backed, having raised about $41M across a 2020 Series A and a 2022 Series B.
People who do well here are comfortable working async in a small, globally distributed team, enjoy owning a wide range of projects, and value flexible hours over fixed office time. Five stated values steer the culture: Bring the weird, Be real with each other, Maintain momentum, Empower others, and Build it together.