Back to The Meridian

Ferro · 6 min read

Automation in the language you're already using

Ferro's project automation is a Rust program, not a Makefile or a folder of shell scripts. Why the xtask pattern pays for itself, and the small details that make automation pleasant to live with.

KR
Karthik Rajkumar28 September 2026 · 6 min read

Every project accumulates tasks that are not building the project. Regenerate the test data. Run all the checks the way the build server runs them. Generate some repetitive code. Format everything.

The traditional home for these is a Makefile, or a folder of shell scripts, or in the Python world a handful of commands in a pyproject.toml. Ferro puts them in a Rust program called xtask, and there is a convention in the Rust community for doing exactly this.

#How it works

xtask is an ordinary binary crate in the workspace, no different structurally from any other program. What makes it feel built-in is three lines in .cargo/config.toml:

.cargo/config.tomlTOML
[alias]
xtask = "run --quiet --package xtask --"

Cargo aliases let you define new subcommands. With that in place, cargo xtask gen-fixtures expands to "build and run the xtask binary, passing it gen-fixtures". It reads like a native Cargo command:

The commandsShell
cargo xtask gen-fixtures     # regenerate the NumPy reference data
cargo xtask check-fixtures   # fail if the committed data is stale
cargo xtask list-fixtures    # show names, types, shapes
cargo xtask ci               # everything the build server runs
cargo xtask miri             # check unsafe code for undefined behaviour

No installation step, no separate tool to keep in sync, nothing to add to your PATH. Anyone who can build the project can run the automation.

#Why not just use shell scripts

Shell scripts are excellent for short things and get progressively worse as they grow. The specific problems that bite:

Portability. Ferro should build on Linux, macOS, and Windows. Shell scripts do not run on Windows without extra machinery, and the versions of sed, find, and date on macOS differ from the GNU ones on Linux in ways that produce silent misbehaviour rather than errors.

No type checking. A typo in a shell variable name is not an error, it is an empty string. That is how rm -rf "$PREFIX/" becomes rm -rf /.

It gets worse with size. Right now the automation shells out to a Python script and runs some Cargo commands, which is exactly what shell is for. But the plan has xtask doing kernel code generation later — writing out the hundreds of near-identical functions that a tensor library needs, one per combination of operation and numeric type. That is real programming with real data structures, and doing it in shell would be miserable.

Writing it in Rust from the start means there is no migration point where the shell script has to be thrown away, and no period where automation lives in two places.

And there is a boring practical benefit: xtask has zero dependencies. It uses only the standard library, so it compiles in about a second from cold. An automation entry point that takes thirty seconds to build before doing anything is an automation entry point people work around.

#The command that matters most

cargo xtask ci runs, in order: formatting check, linter with warnings treated as errors, build, tests, documentation build, a build with only default features, and the fixture drift check.

That list is identical to what the build server runs, and that is the whole point. The failure mode this eliminates is the fifteen-minute round trip: push, wait, get an email about a formatting violation, fix it, push again, wait again. If cargo xtask ci passes locally then the pushed branch will pass too, because they are the same commands in the same order.

The fixture check comes last, deliberately, because it is the only step needing Python — someone without a virtual environment configured still gets useful results from everything above it before hitting the one thing they cannot run.

#Finding Python without making it the user's problem

The fixture generator is a Python script needing NumPy and SciPy. That creates a small, genuinely annoying problem: which Python?

A developer machine can easily have five. The system one that the operating system depends on and you should not install packages into. One from Homebrew. One from a version manager. One inside a project virtual environment. Whichever one python3 happens to resolve to today.

The wrong answer is to run python3 and let it fail. Then the error is a Python traceback ending in ModuleNotFoundError: No module named 'numpy', which does not say which interpreter was used, whether that was the intended one, or what to do about it.

So xtask probes. It tries each candidate in priority order — an explicit override, the project virtual environment, then whatever is on the path — and for each one it checks whether the imports actually work:

Probing each interpreterRust
let probe = Command::new(candidate)
    .args(["-c", "import numpy, scipy"])
    .output();
if matches!(probe, Ok(output) if output.status.success()) {
    return Ok(candidate.clone());
}

The first interpreter that can import both wins. If none can, the error names everything it tried and gives the two commands that fix it:

When no interpreter worksText
xtask: no Python interpreter with NumPy and SciPy found.
Tried: /path/to/.venv/bin/python, python3, python

Set one up with:
  python3 -m venv .venv && .venv/bin/pip install numpy scipy

Or point FERRO_PYTHON at an interpreter that already has them.

This is maybe twenty-five lines of code. It converts the most likely first-run failure from a confusing traceback into an instruction. New contributors hit this exact wall, once, and the difference between those two experiences is out of proportion to the effort.

The probe costs a few milliseconds per candidate, which is invisible next to running the script itself.

#Small things that make the output readable

Two details worth mentioning because they cost nothing and pay off constantly.

Every step announces itself, so a failure in a long run is locatable at a glance:

ProgressText
=== formatting ===
=== clippy ===
=== build ===

And when a command fails, the error includes the command that failed, rendered in full:

A failure you can reproduceText
xtask: `cargo clippy --workspace --all-targets --all-features -- -D warnings`
       failed with exit status: 101

Which means you can copy it, paste it, and reproduce the failure directly rather than guessing what flags the automation used. There is a real category of frustration where a build passes locally but fails in automation, and half of those turn out to be the automation running a slightly different command than the one you ran. Printing it removes the guesswork.

#The thing that surprised me

xtask was written expecting it to be scaffolding — necessary, dull, and unrewarding. It ended up being the most-used piece of Ferro's groundwork by a wide margin, because it is what turns "the milestone is met" from a claim into a command.

The bar for this first milestone was that continuous integration is green and the fixture pipeline works end to end. Without automation, verifying that means remembering seven commands and their flags, and being honest about whether you actually ran all of them. With it, the check is cargo xtask ci and an exit code of zero. That is checkable by anyone, on any machine, without trusting anybody's memory.

Next: the checks that automation runs, and why they were all turned on before there was any code to fail them.

Found this useful? Pass it on.