< Agenda

We Sped Up Apache Commons' CI Pipeline 3x With an Algorithm From 1969

16:10 (30 minutes) · Room 1B · Talk · Intermediate · Engineering

Marjia Siddik
Marjia SiddikPhD Student in Computer Science at Trinity College Dublin
Samson Oloruntola
Samson OloruntolaSoftware Engineer at Fidelity Investments

While working as software engineers, we kept watching our CI pipeline crawl to the finish line; one job would drag on forever while three others finished ages ago and just sat there. We got fed up and decided to fix it. Turns out the answer already existed: Longest Processing Time (LPT), a scheduling algorithm from 1969, decades before GitHub Actions was even a thing. We built TestSplit, a CLI tool that profiles your JUnit test suite and uses LPT to split tests across parallel CI jobs so no single job ends up carrying everyone else. We’ll go through why naive splitting, alphabetical order, and round-robin always end up with one overloaded job and a bunch of idle ones, and why LPT's dead simple rule (sort by duration, fill whichever bucket is emptiest) beats it every time. We also implemented MULTIFIT, which uses binary search to obtain a tighter split when LPT's approximation isn't sufficient. We'll get into the actual problems we ran into building both: tests with ordering dependencies that broke LPT's assumptions, floating-point precision bugs in our binary search, and the challenge of parsing Java source files for dependencies without writing a full AST parser. We ran it against Apache Commons Lang, a real open-source repo with 231 test classes and 64,725 tests, and got a 2.24x speedup on test execution and 3.1x over running it sequentially. We'll also be honest about where it falls short: almost 90% of Commons Lang's tests returned a zero duration in the XML output, which messes with scheduling accuracy. If you've ever sat there watching a CI pipeline crawl and thought "there has to be a better way," this is for you. No ML or academic background needed, just tests that take too long. You'll leave with a real mental model for how a 1969 algorithm still beats most naive approaches to a very modern problem, and a working, open-source CLI tool you can point at your own test suite this week to see the difference yourself.

Marjia Siddik

Marjia Siddik is a PhD student in Computer Science at Trinity College Dublin. She was previously a Software Engineer at Fidelity Investments, building and testing tools used by investment analysts.

Samson Oloruntola

Software Engineer at Fidelity Investments and published researcher in software engineering with work on automated refactoring featured by Springer Nature. Interested in building reliable software and tackling complex engineering problems.