Back to The Meridian

Ferro · 6 min read

Warnings as errors, and a Miri job with nothing to check

Before Ferro had any real code, every check was switched on and set to fail the build — including a tool that hunts for bugs in code the project does not contain yet. An empty codebase is the only time that is painless.

KR
Karthik Rajkumar28 September 2026 · 6 min read

The last piece of Ferro's groundwork was to turn on every check available, set them all to fail the build, and configure a tool that detects bugs in a category of code the project does not contain yet.

Doing this to a codebase with no code in it is the only time it is painless, which is the entire argument.

#The checks

Four kinds of check run on every push. Each protects against a different kind of decay.

Formatting. Rust ships rustfmt, which reformats code to a canonical style. Because there is one official style, formatting is not something anyone needs to have opinions about — it is settled, and CI just verifies that the committed code matches. The benefit is that no code review ever spends a comment on whitespace, and no diff is ever cluttered with reformatting noise that hides the actual change.

The linter. clippy catches things that compile but are probably not what you meant: a loop that would read better as an iterator, a comparison that is always true, a needless allocation. It caught two real things in Ferro's first few hundred lines, one of which was a suggestion that would have introduced a bug (that story is in the .npy reader post).

Tests and documentation. The test suite, plus a check that the documentation actually builds — which catches broken cross-references between modules, a thing that rots invisibly otherwise.

Miri. More on this below, because it is the strange one.

#Warnings are errors

The configuration that matters most is one line:

.github/workflows/ci.ymlYAML
env:
  RUSTFLAGS: -D warnings

-D warnings promotes every compiler warning to a hard error. Unused variable? Build fails. Deprecated function? Build fails. Missing documentation on a public item? Build fails.

This is a policy that people have strong feelings about, and in an established codebase the objections are legitimate. Turn it on somewhere with four hundred existing warnings and you have created a multi-week cleanup that blocks all other work, or — much more likely — a long list of exemptions that gradually becomes indistinguishable from not having it on.

But the dynamics of warning counts are unusual. A codebase with zero warnings tends to stay at zero, because the moment one appears the build breaks and whoever caused it fixes it immediately, while the change is small and they still remember the context. A codebase with four hundred warnings tends to grow to five hundred, because one more is invisible and nobody can tell which ones are new.

Ferro had zero warnings by construction, since it had almost no code. Turning the policy on then cost nothing and locked in the state permanently.

#Every public thing must be documented

One additional lint, on top of the defaults:

Cargo.tomlTOML
[workspace.lints.rust]
missing_docs = "warn"

Combined with warnings-as-errors, this means every public type, function, and enum variant in Ferro must carry a documentation comment or the build fails.

Fourteen crates' worth of documentation sounds like a large tax. At this stage it was around thirty lines, because almost nothing public exists yet. Each empty crate needed a paragraph saying what will live there; the error type needed a line per variant. That is it.

The same requirement introduced a few months from now would mean documenting a tensor type, an autodiff engine, and forty neural network layers, all at once, from memory, long after the decisions were made. Nobody does that well. What actually happens is that the lint gets set back to allow with a note saying "TODO: document things", and that note outlives the project.

There is a secondary effect that turned out to matter more than the documentation itself. Being required to write a sentence about each public item is a mild but real forcing function on API design. An item you cannot describe in one sentence is usually an item doing two things.

#Miri, and why it runs with nothing to check

This is the part that looks most like a waste.

Rust's central promise is memory safety: the compiler proves that your program does not read past the end of a buffer, use memory after freeing it, or race two threads against the same data. This is checked at compile time and it is why Rust exists.

The escape hatch is unsafe. Inside an unsafe block you can do things the compiler cannot verify — raw pointer arithmetic, telling the compiler to reinterpret bytes as a different type, assuming a bounds check is unnecessary. The guarantees do not vanish everywhere, but inside that block the responsibility moves to you.

Ferro will need it. The performance goal is a matrix multiplication within a small factor of a hand-tuned assembly library, and reaching that requires hand-written vector instructions and array accesses whose bounds the compiler cannot prove are safe. That is the legitimate use case, and it will be confined to a few small functions.

The problem with unsafe bugs is that they do not behave like other bugs. A genuine violation is undefined behavior, which means the compiler was permitted to assume it could not happen and may have optimized on that basis. The symptom is not a crash at the mistake. It is a crash somewhere unrelated, or wrong results only in release builds, or correct behavior for six months until a compiler upgrade changes an optimization decision.

Miri is an interpreter that runs your tests while tracking what every pointer is actually allowed to touch, and reports violations precisely where they happen. It is slow — tens to hundreds of times slower than native — which is fine for a test suite and hopeless for real work.

Ferro has no unsafe code at all yet. The .npy reader was deliberately written without it. So the Miri job passes trivially and finds nothing.

It runs anyway, and the reason is about the day it stops passing. Suppose Miri were introduced later, alongside the first SIMD kernel. It reports three violations. Are they from the new kernel, or were they always there in code nobody had checked? You do not know, and finding out means auditing everything at once, in the middle of a performance push, which is the worst possible time.

Introduce it now and it is a green line. When the first unsafe block lands and it goes red, there is exactly one candidate. The value is not in what it finds today — it is in having a known-good baseline, so a future failure is attributable.

#One more job

One more, less philosophical:

.github/workflows/ci.ymlYAML
- name: Default features only
  run: cargo build -p ferro

Ferro's optional components — reinforcement learning, game search, tokenizers, the GPU backend — are behind feature flags, so a default build compiles only the core. This job builds with defaults, while everything else builds with all features on.

Without it, feature gating rots in a specific way. Someone adds code to the neural network layer that calls into the search module, forgetting the dependency is optional. It compiles fine for them, and in every CI job that enables everything. It breaks only for the person who wanted a small build, who gets a confusing error about a missing module and no idea why it works for everybody else.

One extra job, thirty seconds, and the promise that a default build stands on its own stays true.

#The foundations, done

The bar was: continuous integration green, and the fixture pipeline producing files that a Rust test loads and compares. Both hold. cargo xtask ci exits zero with 41 tests passing, no warnings anywhere, and 23 fixtures verified against the generator that produced them.

None of it computes anything. There is no tensor. Nothing in the repository can add two numbers together.

What exists is a shape, a source of truth, a way to report failure, a way to run everything with one command, and a set of checks that are all currently green and will therefore mean something when they stop being green. The actual tensor comes next, and it starts against a harness that has already been tested in the failing direction.

The unglamorous work is done. Next: storage, shapes, and strides — where a tensor library either gets its foundations right or spends six months paying for it.

Found this useful? Pass it on.