The last piece of Ferro's groundwork was to turn on every check available, set them all to fail the build, and configure a tool that detects bugs in a category of code the project does not contain yet.
Doing this to a codebase with no code in it is the only time it is painless, which is the entire argument.
#The checks
Four kinds of check run on every push. Each protects against a different kind of decay.
Formatting. Rust ships rustfmt, which reformats code to a canonical
style. Because there is one official style, formatting is not something anyone
needs to have opinions about — it is settled, and CI just verifies that the
committed code matches. The benefit is that no code review ever spends a
comment on whitespace, and no diff is ever cluttered with reformatting noise
that hides the actual change.
The linter. clippy catches things that compile but are probably not what
you meant: a loop that would read better as an iterator, a comparison that is
always true, a needless allocation. It caught two real things in Ferro's first
few hundred lines, one of which was a suggestion that would have introduced a
bug (that story is in
the .npy reader post).
Tests and documentation. The test suite, plus a check that the documentation actually builds — which catches broken cross-references between modules, a thing that rots invisibly otherwise.
Miri. More on this below, because it is the strange one.
#Warnings are errors
The configuration that matters most is one line:
env:
RUSTFLAGS: -D warnings-D warnings promotes every compiler warning to a hard error. Unused variable?
Build fails. Deprecated function? Build fails. Missing documentation on a
public item? Build fails.
This is a policy that people have strong feelings about, and in an established codebase the objections are legitimate. Turn it on somewhere with four hundred existing warnings and you have created a multi-week cleanup that blocks all other work, or — much more likely — a long list of exemptions that gradually becomes indistinguishable from not having it on.
But the dynamics of warning counts are unusual. A codebase with zero warnings tends to stay at zero, because the moment one appears the build breaks and whoever caused it fixes it immediately, while the change is small and they still remember the context. A codebase with four hundred warnings tends to grow to five hundred, because one more is invisible and nobody can tell which ones are new.
Ferro had zero warnings by construction, since it had almost no code. Turning the policy on then cost nothing and locked in the state permanently.
#Every public thing must be documented
One additional lint, on top of the defaults:
[workspace.lints.rust]
missing_docs = "warn"Combined with warnings-as-errors, this means every public type, function, and enum variant in Ferro must carry a documentation comment or the build fails.
Fourteen crates' worth of documentation sounds like a large tax. At this stage it was around thirty lines, because almost nothing public exists yet. Each empty crate needed a paragraph saying what will live there; the error type needed a line per variant. That is it.
The same requirement introduced a few months from now would mean documenting a
tensor type, an autodiff engine, and forty neural network layers, all at once,
from memory, long after the decisions were made. Nobody does that well. What
actually happens is that the lint gets set back to allow with a note saying
"TODO: document things", and that note outlives the project.
There is a secondary effect that turned out to matter more than the documentation itself. Being required to write a sentence about each public item is a mild but real forcing function on API design. An item you cannot describe in one sentence is usually an item doing two things.
#Miri, and why it runs with nothing to check
This is the part that looks most like a waste.
Rust's central promise is memory safety: the compiler proves that your program does not read past the end of a buffer, use memory after freeing it, or race two threads against the same data. This is checked at compile time and it is why Rust exists.
The escape hatch is unsafe. Inside an unsafe block you can do things the
compiler cannot verify — raw pointer arithmetic, telling the compiler to
reinterpret bytes as a different type, assuming a bounds check is unnecessary.
The guarantees do not vanish everywhere, but inside that block the
responsibility moves to you.
Ferro will need it. The performance goal is a matrix multiplication within a small factor of a hand-tuned assembly library, and reaching that requires hand-written vector instructions and array accesses whose bounds the compiler cannot prove are safe. That is the legitimate use case, and it will be confined to a few small functions.
The problem with unsafe bugs is that they do not behave like other bugs. A
genuine violation is undefined behavior, which means the compiler was
permitted to assume it could not happen and may have optimized on that basis.
The symptom is not a crash at the mistake. It is a crash somewhere unrelated,
or wrong results only in release builds, or correct behavior for six months
until a compiler upgrade changes an optimization decision.
Miri is an interpreter that runs your tests while tracking what every pointer is actually allowed to touch, and reports violations precisely where they happen. It is slow — tens to hundreds of times slower than native — which is fine for a test suite and hopeless for real work.
Ferro has no unsafe code at all yet. The .npy reader was deliberately
written without it. So the Miri job passes trivially and finds nothing.
It runs anyway, and the reason is about the day it stops passing. Suppose Miri were introduced later, alongside the first SIMD kernel. It reports three violations. Are they from the new kernel, or were they always there in code nobody had checked? You do not know, and finding out means auditing everything at once, in the middle of a performance push, which is the worst possible time.
Introduce it now and it is a green line. When the first unsafe block lands and
it goes red, there is exactly one candidate. The value is not in what it finds
today — it is in having a known-good baseline, so a future failure is
attributable.
#One more job
One more, less philosophical:
- name: Default features only
run: cargo build -p ferroFerro's optional components — reinforcement learning, game search, tokenizers, the GPU backend — are behind feature flags, so a default build compiles only the core. This job builds with defaults, while everything else builds with all features on.
Without it, feature gating rots in a specific way. Someone adds code to the neural network layer that calls into the search module, forgetting the dependency is optional. It compiles fine for them, and in every CI job that enables everything. It breaks only for the person who wanted a small build, who gets a confusing error about a missing module and no idea why it works for everybody else.
One extra job, thirty seconds, and the promise that a default build stands on its own stays true.
#The foundations, done
The bar was: continuous integration green, and the fixture pipeline producing
files that a Rust test loads and compares. Both hold. cargo xtask ci exits
zero with 41 tests passing, no warnings anywhere, and 23 fixtures verified
against the generator that produced them.
None of it computes anything. There is no tensor. Nothing in the repository can add two numbers together.
What exists is a shape, a source of truth, a way to report failure, a way to run everything with one command, and a set of checks that are all currently green and will therefore mean something when they stop being green. The actual tensor comes next, and it starts against a harness that has already been tested in the failing direction.
The unglamorous work is done. Next: storage, shapes, and strides — where a tensor library either gets its foundations right or spends six months paying for it.